Curated by

avatar

Stacklist Team

stacklist.com/stacklist-team

More in GitHub repos to check out

See all 34 →

More from Stacklist Team

See all stacks →

LMCache: Supercharge Your LLM with Fast KV Cache Layer

LMCache is a KV cache management layer for scalable LLM inference that converts temporary cache into reusable, persistent knowledge across multiple serving engines and storage backends. It reduces time-to-first-token and improves throughput for long-context and agentic workloads while maintaining vendor neutrality and production-level observability.

View card
Built for AI agentsACO · 1598 tokens

Summary

LMCache is a KV cache management layer for scalable LLM inference that converts temporary cache into reusable, persistent knowledge across multiple serving engines and storage backends. It reduces time-to-first-token and improves throughput for long-context and agentic workloads while maintaining vendor neutrality and production-level observability.

Tags

llm-inference · kv-cache · caching-library · gpu-optimization · distributed-systems · vendor-neutral · observability

Key entities

LMCache (technology, 0.99) · KV Cache (technology, 0.98) · vLLM (technology, 0.92) · Redis (technology, 0.9) · PyTorch (technology, 0.88) · PyTorch Foundation (organization, 0.85) · NVIDIA (organization, 0.9) · CoreWeave (organization, 0.85) · Cohere (organization, 0.85) · CacheBlend (technology, 0.8) · prefix-caching (concept, 0.87) · RAG (concept, 0.85) · GTC 2026 (event, 0.8)

Classification

reference · language en · status final

Provenance

claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 13 Jul 2026