Curated by
More in Local AI & GPUs
See all 24 →More from Stacklist Team
See all stacks →Building a Custom ChatGPT Model from Scratch
AI Engineering project where Alex Lotkov built a 496M-parameter ChatGPT-like model from scratch in 24 hours for $20, achieving 230-340 tokens per second on a 4GB laptop GPU through aggressive memory optimization and custom CUDA kernels. The project demonstrates that model building, memory fitting, training, and efficient serving are distinct engineering challenges, with INT4 quantization providing optimal speed and memory usage.
Built for AI agentsACO · 1058 tokens
Summary
AI Engineering project where Alex Lotkov built a 496M-parameter ChatGPT-like model from scratch in 24 hours for $20, achieving 230-340 tokens per second on a 4GB laptop GPU through aggressive memory optimization and custom CUDA kernels. The project demonstrates that model building, memory fitting, training, and efficient serving are distinct engineering challenges, with INT4 quantization providing optimal speed and memory usage.
Tags
llm-inference · gpu-optimization · cuda-kernels · quantization · ai-engineering · model-training · inference-optimization
Key entities
Alex Lotkov (person, 0.95) · CUDA (technology, 0.95) · FlashAttention (technology, 0.9) · PyTorch (technology, 0.85) · Transformer (technology, 0.95) · INT4 quantization (technology, 0.9) · RTX 3050 (technology, 0.85) · Modal A10 (technology, 0.85) · GPU memory optimization (concept, 0.9) · inference optimization (concept, 0.9) · RAG (concept, 0.85)
Classification
analysis · language en · status final
Provenance
claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 17 Jul 2026