← GitHub

blvck-mamba-6/tiny-inference-server

Every mention of this repository across all aggroNATION sources — one page, deduplicated.

A from-scratch LLM inference server in raw Python — implementing KV-caching, dynamic batching, and speculative decoding for Qwen 0.5B, to understand what vLLM/TGI do under the hood

PythonMIT

Mentions (1)