Curated by

M

Mike Boscia

stacklist.com/michael-boscia-871

More in AI Brains

See all 31 →

More from Mike Boscia

See all stacks →

KV, Prefix, Prompt and Semantic Caching in LLMs

KV, Prefix, Prompt and Semantic Caching in LLMs explains four distinct cache layers in language model inference, their trade-offs, and common problems that inhibit cache reuse. The content covers how these caching mechanisms store different objects, from attention tensors to response strings, and demonstrates their implementation with practical code examples.

View card
Built for AI agentsACO · 5357 tokens

Summary

KV, Prefix, Prompt and Semantic Caching in LLMs explains four distinct cache layers in language model inference, their trade-offs, and common problems that inhibit cache reuse. The content covers how these caching mechanisms store different objects, from attention tensors to response strings, and demonstrates their implementation with practical code examples.

Tags

kv-cache · llm-optimization · prompt-caching · semantic-cache · inference · memory-bandwidth · attention-mechanism

Key entities

Avi Chawla (person, 0.95) · KV cache (technology, 0.98) · Prefix caching (technology, 0.95) · Prompt caching (technology, 0.95) · Semantic cache (technology, 0.95) · transformers library (technology, 0.92) · DynamicCache (technology, 0.9) · Anthropic (organization, 0.88) · sentence-transformers (technology, 0.85) · attention-tensors (concept, 0.92) · causal-masking (concept, 0.9) · memory-bandwidth (concept, 0.88)

Classification

tutorial · language en · status final

Provenance

claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 11 Sep 2026