← Back to home

HuggingFace

300 items in the index — today's curated AI papers from HuggingFace Daily. Everything opens right here on the site — you never leave.

Index updated 1h ago · refreshes hourly

← HighlightsStrict order
HuggingFace Daily Papers

Modality-Autoregressive World-Action Models

World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images. Other visual modalities such as depth, pretrained visual features, and point tracks can more efficiently capture geometric, semantic, and motion features. Howe

Adam Hung, Bardienus P. Duisterhof, Deva Ramanan, Jeffrey Ichnowski · Sep 15, 202634
7 likes
HuggingFace Daily Papers

PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control

Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet existing controllable methods either require the full control schedule before generation starts, or use pixel-space signals that d

Chuhao Chen, Peter Wonka, Chaoyang Wang, Chen Wang · Sep 15, 202634
3 likes
HuggingFace Daily Papers

FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This setup bottlenecks generative performance behind frozen embeddings. To bridge th

Guangyu Sun, Shlok Kumar Mishra, Wentao Bao, Robert Zhenheng Yang · Sep 15, 202634
12 likes