← Back to home

HuggingFace

307 items in the index — today's curated AI papers from HuggingFace Daily. Everything opens right here on the site — you never leave.

Index updated 1m ago · refreshes hourly

← HighlightsStrict order
HuggingFace Daily Papers

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision coreof household robots. Existing evaluations,

Yizhan Li, Jianxin You, Mengyang Xiong, Yinhuan Chen · Sep 9, 202624
41 likes
HuggingFace Daily Papers

Towards a Deterministic Math Solver for Clinical Language Models

Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard response is to hardcode each calculator as a validated function, one at a time. We test an alternative: the mode

Felipe Ocampo Osorio, Sebastián Andrés Cajas Ordoñez, Maximin Lange, Rafi Al Attrach · Sep 9, 202624
2 likes
HuggingFace Daily Papers

NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducin

NCP Team, Jiaqi Cao, Chiyu Chen, Shuang Cheng · Sep 9, 202626
304 likes
HuggingFace Daily Papers

Programmable World Model

Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent world state and enforcing programmable rules over extended interactions. We introduce Programmable World Model, a framework that d

Zheng-Hui Huang, Guixu Lin, Jiacheng Lin, Yi-Chuan Huang · Sep 9, 202633
84 likes
HuggingFace Daily Papers

Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning

Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain overwhelmingly English-centric: models primarily reason in English regardless of the language they are prompted in. This is inaccessible for non-English-speaking

Mehrnaz Mofakhami, Ananya Sahu, Alejandro R. Salamanca, Daniel D'souza · Sep 9, 202626
28 likes
HuggingFace Daily Papers

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrie

Killian Steunou, Yannis Tevissen, Mounîm A. El Yacoubi · Sep 9, 202627
13 likes
HuggingFace Daily Papers

Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States

Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated text. We propose a reference-based method that audits bias in hidden-state representations across rel

Marek Jeliński, Jan Dubiński, Maciej Chrabaszcz, Sebastian Cygert · Sep 9, 202627
5 likes