Curated by
More in Local AI & GPUs
See all 24 →More from Stacklist Team
See all stacks →Run GLM-5.2 on a Consumer Machine with Colibri
Colibri is a lightweight C-based runtime framework for running the 744B-parameter GLM-5.2 Mixture-of-Experts model on consumer machines with ~25GB RAM through disk-based expert streaming and efficient memory management. The framework implements faithful GLM-5.2 forward passes with MLA attention, native MTP speculative decoding, and grammar-forced decoding for constrained outputs.
Built for AI agentsACO · 10583 tokens
Summary
Colibri is a lightweight C-based runtime framework for running the 744B-parameter GLM-5.2 Mixture-of-Experts model on consumer machines with ~25GB RAM through disk-based expert streaming and efficient memory management. The framework implements faithful GLM-5.2 forward passes with MLA attention, native MTP speculative decoding, and grammar-forced decoding for constrained outputs.
Tags
moe-runtime · lightweight-framework · glm-5.2 · speculative-decoding · quantization · expert-streaming
Key entities
Colibri (technology, 0.99) · GLM-5.2 (technology, 0.99) · Mixture-of-Experts (concept, 0.95) · MLA attention (technology, 0.9) · speculative-decoding (concept, 0.92) · CUDA (technology, 0.88) · int4-quantization (concept, 0.9) · DeepSeek-V3 (technology, 0.85)
Classification
framework · language en · status final
Provenance
claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 13 Jul 2026