← Back to home

#huggingface

303 items tagged “huggingface” across every source.

HuggingFace Daily Papers

Modality-Autoregressive World-Action Models

World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images. Other visual modalities such as depth, pretrained visual features, and point tracks can more efficiently capture geometric, semantic, and motion features. Howe

Adam Hung, Bardienus P. Duisterhof, Deva Ramanan, Jeffrey Ichnowski · Sep 15, 202634
7 likes
HuggingFace Daily Papers

PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control

Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet existing controllable methods either require the full control schedule before generation starts, or use pixel-space signals that d

Chuhao Chen, Peter Wonka, Chaoyang Wang, Chen Wang · Sep 15, 202634
3 likes
HuggingFace Daily Papers

FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This setup bottlenecks generative performance behind frozen embeddings. To bridge th

Guangyu Sun, Shlok Kumar Mishra, Wentao Bao, Robert Zhenheng Yang · Sep 15, 202634
12 likes
HuggingFace Daily Papers

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from

Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li · Sep 14, 202632
1 likes
HuggingFace Daily Papers

Omni-Streaming Thinking

Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If that interpretation enters memory as a fact, later reasonin

Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li · Sep 14, 202634
19 likes