Swiftlet: streams 80B Qwen MoE experts from iPhone SSD, 2.5GB RAM, one token per second
Github Awesome · Aug 4, 2026 · 4,034 views · 89 likes · 3 comments on YouTube
Description
Swiftlet runs 35B and 80B Qwen mixture-of-experts models on Apple devices by keeping the dense core in memory and streaming routed experts from SSD. QPack files turn each expert fetch into one read, while caching and Metal kernels handle inference. The maintainer reports the 35B
Comments
Sign in to join the discussion.