{"version":"1.0","type":"card","id":"f02e1fa8-7830-4bc3-8d25-ea3034079f45","url":"https://stacklist.com/card/f02e1fa8-7830-4bc3-8d25-ea3034079f45","title":"25x inference performance on NVIDIA GB300 NVL72","source_url":"https://www.linkedin.com/posts/paoloperrone_25x-inference-performance-on-nvidia-gb300-share-7478555102983827456-9R0I/?utm_source=share&utm_medium=member_ios&rcm=ACoAAAI21ZsBNnZPaKuTab7nquKLCveUW7o-1DE","note":"This page discusses the significant improvements in inference performance achieved with the NVIDIA GB300 NVL72 using the SGLang framework and RadixAttention mechanism. It highlights how this approach minimizes recomputation of key-value caches, resulting in a 25x throughput increase for agent workloads.","image":{"url":"https://ucarecdn.com/74e0aae0-89b2-4355-a7f3-997896f89b2b/","alt":"25x inference performance on NVIDIA GB300 NVL72","width":1280,"height":800},"stack":{"id":"4aae218c-38d7-4c05-b7ae-f2db0029a2e8","title":"Local AI & GPUs","url":"https://stacklist.com/stack/4aae218c-38d7-4c05-b7ae-f2db0029a2e8"},"created_at":"2026-07-06T01:23:16.617Z","updated_at":null,"aco":{"summary":"25x inference performance improvement on NVIDIA GB300 achieved through SGLang's RadixAttention mechanism, which optimizes KV cache reuse by indexing prefixes in a radix tree to eliminate redundant computations. For agent workloads with 80% prefix overlap, this approach combined with chunked prefill and continuous batching delivers significant throughput gains by computing shared system prompts and context only once instead of repeatedly.","tags":["inference-optimization","radix-attention","kv-cache","sglang","agent-workloads","serving-framework","gpu-performance"],"key_entities":[{"name":"Paolo Perrone","type":"person","confidence":0.95},{"name":"NVIDIA GB300","type":"technology","confidence":0.95},{"name":"SGLang","type":"technology","confidence":0.95},{"name":"RadixAttention","type":"technology","confidence":0.95},{"name":"KV cache","type":"concept","confidence":0.95},{"name":"prefix-matching","type":"concept","confidence":0.9},{"name":"xAI","type":"organization","confidence":0.85},{"name":"NVIDIA","type":"organization","confidence":0.9},{"name":"AMD","type":"organization","confidence":0.85},{"name":"Microsoft Azure","type":"organization","confidence":0.85},{"name":"AWS","type":"organization","confidence":0.85},{"name":"Cursor","type":"organization","confidence":0.8},{"name":"agent-workloads","type":"concept","confidence":0.9},{"name":"Grok","type":"technology","confidence":0.85}],"classification":"analysis","language":"en","confidence":0.85,"provenance":{"model":"claude-haiku-4-5","tool":"@stacklist/be@0.1.0","confidence":0.85,"timestamp":"2026-07-06T01:23:22.208Z"},"token_counts":{"approximate":553,"cl100k":492},"content_hash":"sha256:98450f2f153065f6eb8e24d868e713ef6f53f6bff642f0cd97aec8fd0c4168e5","acp_version":"0.2","body_available":true,"body_tokens":553,"visibility":"public","agent_accessible":true,"status":"final"},"_links":{"self":"/api/public/card/f02e1fa8-7830-4bc3-8d25-ea3034079f45.json","html":"https://stacklist.com/card/f02e1fa8-7830-4bc3-8d25-ea3034079f45","md":"/api/public/card/f02e1fa8-7830-4bc3-8d25-ea3034079f45.md","stack_json":"/api/public/stack/4aae218c-38d7-4c05-b7ae-f2db0029a2e8.json"}}