Curated by
M
More from Mike Boscia
See all stacks →Everything You Need To Know About Inference Engines
This page provides an overview of inference engines and their role in running large language models (LLMs) locally at home. It discusses key concepts such as the differences between prefill and decode, VRAM and bandwidth, and the significance of KV cache and quantization.
Built for AI agentsNo ACO on this card