Curated by

M

Mike Boscia

stacklist.com/michael-boscia-871

More in Local AI & GPUs

See all 36 →

More from Mike Boscia

See all stacks →

UC Berkeley Open-Sources FreeToken for LLM Inference

FreeToken is an open-source inference engine from UC Berkeley that achieves 2-4x faster local LLM inference than Ollama by intelligently routing Mixture-of-Experts models between GPU and CPU memory. The system dynamically profiles hardware bandwidth and optimizes expert placement, enabling large models like Qwen 35B to run on consumer 8GB GPUs.

View card
Built for AI agentsACO · 893 tokens

Summary

FreeToken is an open-source inference engine from UC Berkeley that achieves 2-4x faster local LLM inference than Ollama by intelligently routing Mixture-of-Experts models between GPU and CPU memory. The system dynamically profiles hardware bandwidth and optimizes expert placement, enabling large models like Qwen 35B to run on consumer 8GB GPUs.

Tags

freetoken · llm-inference · mixture-of-experts · gpu-optimization · open-source · local-inference

Key entities

FreeToken (technology, 0.99) · UC Berkeley (organization, 0.99) · Akshay Pachaar (person, 0.99) · Qwen3.6-35B (technology, 0.95) · DeepSeek-V4-Flash (technology, 0.95) · GLM-5.2 (technology, 0.95) · Ollama (technology, 0.95) · Mixture-of-Experts (concept, 0.98) · OpenAI API (technology, 0.9) · Anthropic API (technology, 0.9)

Classification

analysis · language en · status final

Provenance

claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 24 Aug 2026