← Back to home

HuggingFace

300 items in the index — today's curated AI papers from HuggingFace Daily. Everything opens right here on the site — you never leave.

Index updated 2h ago · refreshes hourly

← HighlightsStrict order
HuggingFace Daily Papers

StepAudio 3 Realtime Technical Report

Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop. Deep Perception captures rich acoustic cues to interpret

Bin Lin, Bo Zhao, Boyang Zhang, Boyong Wu · Sep 12, 202627
99 likes
HuggingFace Daily Papers

StepAudio 3 Music Technical Report

We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation. The StepAudio Music Tokenizer represents audio as a 50-Hz stream from a 65536-entry single codebook, using semantica

Chengli Feng, Zhiyue Wu, Jiahao Song, Zheqi Dai · Sep 11, 202626
75 likes