Curated by
More in GitHub repos to check out
See all 34 →More from Stacklist Team
See all stacks →LMCache: Supercharge Your LLM with Fast KV Cache Layer
LMCache is a KV cache management layer for scalable LLM inference that converts temporary cache into reusable, persistent knowledge across multiple serving engines and storage backends. It reduces time-to-first-token and improves throughput for long-context and agentic workloads while maintaining vendor neutrality and production-level observability.
Built for AI agentsACO · 1598 tokens
Summary
LMCache is a KV cache management layer for scalable LLM inference that converts temporary cache into reusable, persistent knowledge across multiple serving engines and storage backends. It reduces time-to-first-token and improves throughput for long-context and agentic workloads while maintaining vendor neutrality and production-level observability.
Tags
llm-inference · kv-cache · caching-library · gpu-optimization · distributed-systems · vendor-neutral · observability
Key entities
LMCache (technology, 0.99) · KV Cache (technology, 0.98) · vLLM (technology, 0.92) · Redis (technology, 0.9) · PyTorch (technology, 0.88) · PyTorch Foundation (organization, 0.85) · NVIDIA (organization, 0.9) · CoreWeave (organization, 0.85) · Cohere (organization, 0.85) · CacheBlend (technology, 0.8) · prefix-caching (concept, 0.87) · RAG (concept, 0.85) · GTC 2026 (event, 0.8)
Classification
reference · language en · status final
Provenance
claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 13 Jul 2026