human-agent-society/reef_
Continual learning infra for self-improving agents
7 items tagged “inference” across every source.
Continual learning infra for self-improving agents
CROW - Your AI Agent. MCPs, OpenRouter, Any Model or local. It's your choice.
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang is a high-performance serving framework for large language models and multimodal models.
Reproducible Qwen3.8-27B serving and native rank-1 LoRA experiments on NVIDIA DGX Spark
GLM-5.3-Flash on two DGX Sparks: opinionated vLLM/EXL3 serving, tinyGLM gates, and reproducible inference research.
A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability