REPO //Trendshift
blvck-mamba-6/tiny-inference-server_
A from-scratch LLM inference server in raw Python — implementing KV-caching, dynamic batching, and speculative decoding for Qwen 0.5B, to understand what vLLM/TGI do under the hood
★ 0⑂ 0Python
Sep 7, 2026[40]
https://github.com/blvck-mamba-6/tiny-inference-server