{"version":"1.0","type":"card","id":"6da51950-642f-4126-bce2-bbc36467971a","url":"https://stacklist.com/card/6da51950-642f-4126-bce2-bbc36467971a","title":"Run GLM-5.2 on a Consumer Machine with Colibri","source_url":"https://github.com/JustVugg/colibri","note":"Colibri allows you to run the GLM-5.2 model (744B MoE) on a consumer machine with 25GB of RAM using pure C and zero dependencies. This tiny engine streams experts from disk, making it efficient and powerful for various applications.","image":{"url":"https://ucarecdn.com/4876039e-4341-4709-ae35-cefc332eee1a/","alt":"Run GLM-5.2 on a Consumer Machine with Colibri","width":1200,"height":600},"stack":{"id":"4aae218c-38d7-4c05-b7ae-f2db0029a2e8","title":"Local AI & GPUs","url":"https://stacklist.com/stack/4aae218c-38d7-4c05-b7ae-f2db0029a2e8"},"created_at":"2026-07-13T22:30:10.220Z","updated_at":null,"aco":{"summary":"Colibri is a lightweight C-based runtime framework for running the 744B-parameter GLM-5.2 Mixture-of-Experts model on consumer machines with ~25GB RAM through disk-based expert streaming and efficient memory management. The framework implements faithful GLM-5.2 forward passes with MLA attention, native MTP speculative decoding, and grammar-forced decoding for constrained outputs.","tags":["moe-runtime","lightweight-framework","glm-5.2","speculative-decoding","quantization","expert-streaming"],"key_entities":[{"name":"Colibri","type":"technology","confidence":0.99},{"name":"GLM-5.2","type":"technology","confidence":0.99},{"name":"Mixture-of-Experts","type":"concept","confidence":0.95},{"name":"MLA attention","type":"technology","confidence":0.9},{"name":"speculative-decoding","type":"concept","confidence":0.92},{"name":"CUDA","type":"technology","confidence":0.88},{"name":"int4-quantization","type":"concept","confidence":0.9},{"name":"DeepSeek-V3","type":"technology","confidence":0.85}],"classification":"framework","language":"en","confidence":0.85,"provenance":{"model":"claude-haiku-4-5","tool":"@stacklist/be@0.1.0","confidence":0.85,"timestamp":"2026-07-13T22:30:20.451Z"},"token_counts":{"approximate":10583,"cl100k":11154},"content_hash":"sha256:ee595c0e4b5a00456c30ddf7084a4be3d30fba9ac57ce7c81b82a619e2cf8509","acp_version":"0.2","body_available":true,"body_tokens":10583,"visibility":"public","agent_accessible":true,"status":"final"},"_links":{"self":"/api/public/card/6da51950-642f-4126-bce2-bbc36467971a.json","html":"https://stacklist.com/card/6da51950-642f-4126-bce2-bbc36467971a","md":"/api/public/card/6da51950-642f-4126-bce2-bbc36467971a.md","stack_json":"/api/public/stack/4aae218c-38d7-4c05-b7ae-f2db0029a2e8.json"}}