Curated by

avatar

Stacklist Team

stacklist.com/stacklist-team

More in Local AI & GPUs

See all 24 →

More from Stacklist Team

See all stacks →

LLM Engineer's Almanac - Advisor | Modal

Llama 3.1 8B Model Advisor is an interactive benchmarking tool that measures per-replica throughput and client-side latency for open-weight language models running on inference engines like vLLM, SGLang, and TensorRT-LLM. Users can select from various models and workload configurations to compare performance metrics and identify optimal deployment strategies.

View card
Built for AI agentsACO · 2066 tokens

Summary

Llama 3.1 8B Model Advisor is an interactive benchmarking tool that measures per-replica throughput and client-side latency for open-weight language models running on inference engines like vLLM, SGLang, and TensorRT-LLM. Users can select from various models and workload configurations to compare performance metrics and identify optimal deployment strategies.

Tags

llm-benchmarking · inference-engines · model-performance · latency-throughput · quantization · open-source

Key entities

Llama 3.1 8B (technology, 0.95) · vLLM (technology, 0.92) · SGLang (technology, 0.92) · TensorRT-LLM (technology, 0.92) · Qwen 2.5 7B (technology, 0.88) · DeepSeek-V3 (technology, 0.88) · Modal (organization, 0.9) · Gemma 3 (technology, 0.85) · quantization-formats (concept, 0.87) · time-to-first-token (concept, 0.85)

Classification

reference · language en · status final

Provenance

claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 18 Jul 2026