Curated by
More in GitHub repos to check out
See all 34 →More from Stacklist Team
See all stacks →Runs 405B LLMs on 8GB VRAM
AirLLM is an AI toolkit that optimizes inference memory usage, enabling large language models like 70B to run on single 4GB GPUs without quantization or pruning. It supports multiple model architectures and includes model compression features for up to 3x inference speed improvement.
Built for AI agentsACO · 2292 tokens
Summary
AirLLM is an AI toolkit that optimizes inference memory usage, enabling large language models like 70B to run on single 4GB GPUs without quantization or pruning. It supports multiple model architectures and includes model compression features for up to 3x inference speed improvement.
Tags
large-language-models · memory-optimization · gpu-inference · model-compression · quantization · open-source · toolkit
Key entities
AirLLM (technology, 0.99) · Llama3.1 (technology, 0.95) · Qwen2.5 (technology, 0.9) · Llama3 (technology, 0.95) · ChatGLM (technology, 0.85) · Mistral (technology, 0.85) · bitsandbytes (technology, 0.8) · block-wise quantization (concept, 0.85) · model compression (concept, 0.9) · Hugging Face (organization, 0.85)
Classification
framework · language en · status final
Provenance
claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 17 Jun 2026