Curated by

avatar

Stacklist Team

stacklist.com/stacklist-team

More in GitHub repos to check out

See all 34 →

More from Stacklist Team

See all stacks →

Runs 405B LLMs on 8GB VRAM

AirLLM is an AI toolkit that optimizes inference memory usage, enabling large language models like 70B to run on single 4GB GPUs without quantization or pruning. It supports multiple model architectures and includes model compression features for up to 3x inference speed improvement.

View card
Built for AI agentsACO · 2292 tokens

Summary

AirLLM is an AI toolkit that optimizes inference memory usage, enabling large language models like 70B to run on single 4GB GPUs without quantization or pruning. It supports multiple model architectures and includes model compression features for up to 3x inference speed improvement.

Tags

large-language-models · memory-optimization · gpu-inference · model-compression · quantization · open-source · toolkit

Key entities

AirLLM (technology, 0.99) · Llama3.1 (technology, 0.95) · Qwen2.5 (technology, 0.9) · Llama3 (technology, 0.95) · ChatGLM (technology, 0.85) · Mistral (technology, 0.85) · bitsandbytes (technology, 0.8) · block-wise quantization (concept, 0.85) · model compression (concept, 0.9) · Hugging Face (organization, 0.85)

Classification

framework · language en · status final

Provenance

claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 17 Jun 2026