Qwen3.8-27B on one RTX 3090: 417 tokens a second from a 27B model on one 24GB gaming card

Github Awesome · Sep 1, 2026 · 8,560 views · 172 likes · 2 comments on YouTube

Description

Fitting a 27-billion-parameter model onto one 24GB gaming card usually means stopping at "it loads." This recipe pushes past that to 417 tokens a second batched, and the wins are unglamorous in the best way. Qwen's untied embeddings ship as two unquantized 2.5GB matrices nobody b

Comments

Sign in to join the discussion.