Curated by

avatar

Stacklist Team

stacklist.com/stacklist-team

More in Local AI & GPUs

See all 24 →

More from Stacklist Team

See all stacks →

Building a Custom ChatGPT Model from Scratch

AI Engineering project where Alex Lotkov built a 496M-parameter ChatGPT-like model from scratch in 24 hours for $20, achieving 230-340 tokens per second on a 4GB laptop GPU through aggressive memory optimization and custom CUDA kernels. The project demonstrates that model building, memory fitting, training, and efficient serving are distinct engineering challenges, with INT4 quantization providing optimal speed and memory usage.

View card
Built for AI agentsACO · 1058 tokens

Summary

AI Engineering project where Alex Lotkov built a 496M-parameter ChatGPT-like model from scratch in 24 hours for $20, achieving 230-340 tokens per second on a 4GB laptop GPU through aggressive memory optimization and custom CUDA kernels. The project demonstrates that model building, memory fitting, training, and efficient serving are distinct engineering challenges, with INT4 quantization providing optimal speed and memory usage.

Tags

llm-inference · gpu-optimization · cuda-kernels · quantization · ai-engineering · model-training · inference-optimization

Key entities

Alex Lotkov (person, 0.95) · CUDA (technology, 0.95) · FlashAttention (technology, 0.9) · PyTorch (technology, 0.85) · Transformer (technology, 0.95) · INT4 quantization (technology, 0.9) · RTX 3050 (technology, 0.85) · Modal A10 (technology, 0.85) · GPU memory optimization (concept, 0.9) · inference optimization (concept, 0.9) · RAG (concept, 0.85)

Classification

analysis · language en · status final

Provenance

claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 17 Jul 2026