Swiftlet: streams 80B Qwen MoE experts from iPhone SSD, 2.5GB RAM, one token per second

Github Awesome · Aug 4, 2026 · 4,034 views · 89 likes · 3 comments on YouTube

Description

Swiftlet runs 35B and 80B Qwen mixture-of-experts models on Apple devices by keeping the dense core in memory and streaming routed experts from SSD. QPack files turn each expert fetch into one read, while caching and Metal kernels handle inference. The maintainer reports the 35B

Comments

Sign in to join the discussion.