{"version":"1.0","type":"card","id":"2a967f43-200e-4e9a-8e39-a14d3101ec41","url":"https://stacklist.com/card/2a967f43-200e-4e9a-8e39-a14d3101ec41","title":"Everything You Need To Know About Inference Engines","source_url":"https://x.com/TheAhmadOsman/status/2062646043687166225","note":"This page provides an overview of inference engines and their role in running large language models (LLMs) locally at home. It discusses key concepts such as the differences between prefill and decode, VRAM and bandwidth, and the significance of KV cache and quantization.","image":{"url":"https://ucarecdn.com/32e93026-5443-456c-888b-fa324e81c740/","alt":"Everything You Need To Know About Inference Engines","width":1200,"height":1200},"stack":{"id":"4d7ef7fd-5f42-4150-adcc-f1e3197f7cbf","title":"X Bookmarks","url":"https://stacklist.com/c/technology/stack/4d7ef7fd-5f42-4150-adcc-f1e3197f7cbf"},"created_at":"2026-07-29T02:53:30.879Z","updated_at":null,"aco":null,"_links":{"self":"/api/public/card/2a967f43-200e-4e9a-8e39-a14d3101ec41.json","html":"https://stacklist.com/card/2a967f43-200e-4e9a-8e39-a14d3101ec41","md":"/api/public/card/2a967f43-200e-4e9a-8e39-a14d3101ec41.md","stack_json":"/api/public/stack/4d7ef7fd-5f42-4150-adcc-f1e3197f7cbf.json"}}