Curated by
More in Local AI & GPUs
See all 24 →More from Stacklist Team
See all stacks →CUDA-LLM: A GitHub Repository for Development
CUDA LLM is a GPU-optimized decoder-only language model implementation featuring custom C++/CUDA attention kernels, INT8/INT4 quantization, and a complete training and serving stack designed to run on laptop GPUs. The 496M parameter model achieves 230 tok/s on laptop INT4 decode and includes advanced features like sparse attention, RAG integration, and production-grade deployment capabilities.
Built for AI agentsACO · 9428 tokens
Summary
CUDA LLM is a GPU-optimized decoder-only language model implementation featuring custom C++/CUDA attention kernels, INT8/INT4 quantization, and a complete training and serving stack designed to run on laptop GPUs. The 496M parameter model achieves 230 tok/s on laptop INT4 decode and includes advanced features like sparse attention, RAG integration, and production-grade deployment capabilities.
Tags
cuda · llm · gpu-optimization · transformer · quantization · inference · flashattention
Key entities
CUDA (technology, 0.99) · PyTorch (technology, 0.95) · FlashAttention (technology, 0.92) · HuggingFace (technology, 0.88) · Modal (technology, 0.9) · Safetensors (technology, 0.85) · RoPE (technology, 0.88) · RMSNorm (technology, 0.85) · SwiGLU (technology, 0.85) · BPE tokenizer (technology, 0.87) · LLaMA (technology, 0.9) · quantization (concept, 0.95) · sparse-attention (concept, 0.92) · RAG (concept, 0.9) · instruction-fine-tuning (concept, 0.88) · laptop-gpu (location, 0.85)
Classification
reference · language en · status final
Provenance
claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 17 Jul 2026