Curated by

M

Mike Boscia

stacklist.com/michael-boscia-871

More in Local AI & GPUs

See all 36 →

More from Mike Boscia

See all stacks →

25x inference performance on NVIDIA GB300 NVL72

This page discusses the significant improvements in inference performance achieved with the NVIDIA GB300 NVL72 using the SGLang framework and RadixAttention mechanism. It highlights how this approach minimizes recomputation of key-value caches, resulting in a 25x throughput increase for agent workloads.

View card
Built for AI agentsNo ACO on this card