Curated by
More in Local AI & GPUs
See all 36 →More from Mike Boscia
See all stacks →UC Berkeley Open-Sources FreeToken for LLM Inference
FreeToken is an open-source inference engine from UC Berkeley that achieves 2-4x faster local LLM inference than Ollama by intelligently routing Mixture-of-Experts models between GPU and CPU memory. The system dynamically profiles hardware bandwidth and optimizes expert placement, enabling large models like Qwen 35B to run on consumer 8GB GPUs.
Built for AI agentsACO · 893 tokens
Summary
FreeToken is an open-source inference engine from UC Berkeley that achieves 2-4x faster local LLM inference than Ollama by intelligently routing Mixture-of-Experts models between GPU and CPU memory. The system dynamically profiles hardware bandwidth and optimizes expert placement, enabling large models like Qwen 35B to run on consumer 8GB GPUs.
Tags
freetoken · llm-inference · mixture-of-experts · gpu-optimization · open-source · local-inference
Key entities
FreeToken (technology, 0.99) · UC Berkeley (organization, 0.99) · Akshay Pachaar (person, 0.99) · Qwen3.6-35B (technology, 0.95) · DeepSeek-V4-Flash (technology, 0.95) · GLM-5.2 (technology, 0.95) · Ollama (technology, 0.95) · Mixture-of-Experts (concept, 0.98) · OpenAI API (technology, 0.9) · Anthropic API (technology, 0.9)
Classification
analysis · language en · status final
Provenance
claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 24 Aug 2026