khimaros/aimbot_
ai model quality research
6 items tagged “llama-cpp” across every source.
ai model quality research
CROW - Your AI Agent. MCPs, OpenRouter, Any Model or local. It's your choice.
Share the model on your machine with friends: one binary in front of llama.cpp, vLLM, Ollama or LM Studio; one invite code; they chat from a browser. Self-hosted, end-to-end encrypted, no accounts.
A tiered-memory system design for workloads that don't fit in RAM: measure the working set, pin the hot tier, stream the cold tier from flash. Ships the residency calculator, measurement harnesses, and the build recipes behind it. Predictions validated against public benchmarks.
Apple-native self-hosting CLI: OCI containers via Apple's container runtime + Metal-backed llama.cpp serving, with honest unified-memory reporting
LLM speculative inference server for heterogeneous hardware & consumer GPUs