Curated by

avatar

Stacklist Team

stacklist.com/stacklist-team

More in Local AI & GPUs

See all 24 →

More from Stacklist Team

See all stacks →

Run GLM-5.2 on a Consumer Machine with Colibri

Colibri is a lightweight C-based runtime framework for running the 744B-parameter GLM-5.2 Mixture-of-Experts model on consumer machines with ~25GB RAM through disk-based expert streaming and efficient memory management. The framework implements faithful GLM-5.2 forward passes with MLA attention, native MTP speculative decoding, and grammar-forced decoding for constrained outputs.

View card
Built for AI agentsACO · 10583 tokens

Summary

Colibri is a lightweight C-based runtime framework for running the 744B-parameter GLM-5.2 Mixture-of-Experts model on consumer machines with ~25GB RAM through disk-based expert streaming and efficient memory management. The framework implements faithful GLM-5.2 forward passes with MLA attention, native MTP speculative decoding, and grammar-forced decoding for constrained outputs.

Tags

moe-runtime · lightweight-framework · glm-5.2 · speculative-decoding · quantization · expert-streaming

Key entities

Colibri (technology, 0.99) · GLM-5.2 (technology, 0.99) · Mixture-of-Experts (concept, 0.95) · MLA attention (technology, 0.9) · speculative-decoding (concept, 0.92) · CUDA (technology, 0.88) · int4-quantization (concept, 0.9) · DeepSeek-V3 (technology, 0.85)

Classification

framework · language en · status final

Provenance

claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 13 Jul 2026