← Back to home

HuggingFace

300 items in the index — today's curated AI papers from HuggingFace Daily. Everything opens right here on the site — you never leave.

Index updated 1h ago · refreshes hourly

← HighlightsStrict order
HuggingFace Daily Papers

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from

Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li · Sep 14, 202632
1 likes
HuggingFace Daily Papers

Omni-Streaming Thinking

Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If that interpretation enters memory as a fact, later reasonin

Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li · Sep 14, 202634
19 likes
HuggingFace Daily Papers

Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training

Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current policy: some are already saturated while others are too difficult to yield reliable learning signals, yet both receive equal rollout budget under standard train

Yuanhao Yue, Qianli Ma, Chengyu Wang, Haoting Wang · Sep 14, 202633
9 likes