{"version":"1.0","type":"card","id":"62524f91-af2a-4b8d-885a-412c7591bacb","url":"https://stacklist.com/card/62524f91-af2a-4b8d-885a-412c7591bacb","title":"25x inference performance on NVIDIA GB300 NVL72","source_url":"https://www.linkedin.com/posts/paoloperrone_25x-inference-performance-on-nvidia-gb300-share-7478555102983827456-9R0I/?utm_source=share&utm_medium=member_ios&rcm=ACoAAAI21ZsBNnZPaKuTab7nquKLCveUW7o-1DE","note":"This page discusses the significant improvements in inference performance achieved with the NVIDIA GB300 NVL72 using the SGLang framework and RadixAttention mechanism. It highlights how this approach minimizes recomputation of key-value caches, resulting in a 25x throughput increase for agent workloads.","image":{"url":"https://ucarecdn.com/883695fc-f0de-4753-92e9-1a7ddc6afa45/","alt":"25x inference performance on NVIDIA GB300 NVL72","width":1200,"height":750},"stack":{"id":"ed01fcda-d7b5-4a4b-bdfa-bc07c9df5228","title":"Local AI & GPUs","url":"https://stacklist.com/c/technology/stack/ed01fcda-d7b5-4a4b-bdfa-bc07c9df5228"},"created_at":"2026-07-29T02:39:11.042Z","updated_at":null,"aco":null,"_links":{"self":"/api/public/card/62524f91-af2a-4b8d-885a-412c7591bacb.json","html":"https://stacklist.com/card/62524f91-af2a-4b8d-885a-412c7591bacb","md":"/api/public/card/62524f91-af2a-4b8d-885a-412c7591bacb.md","stack_json":"/api/public/stack/ed01fcda-d7b5-4a4b-bdfa-bc07c9df5228.json"}}