Curated by

avatar

Stacklist Team

stacklist.com/stacklist-team

More in Local AI & GPUs

See all 24 →

More from Stacklist Team

See all stacks →

CUDA-LLM: A GitHub Repository for Development

CUDA LLM is a GPU-optimized decoder-only language model implementation featuring custom C++/CUDA attention kernels, INT8/INT4 quantization, and a complete training and serving stack designed to run on laptop GPUs. The 496M parameter model achieves 230 tok/s on laptop INT4 decode and includes advanced features like sparse attention, RAG integration, and production-grade deployment capabilities.

View card
Built for AI agentsACO · 9428 tokens

Summary

CUDA LLM is a GPU-optimized decoder-only language model implementation featuring custom C++/CUDA attention kernels, INT8/INT4 quantization, and a complete training and serving stack designed to run on laptop GPUs. The 496M parameter model achieves 230 tok/s on laptop INT4 decode and includes advanced features like sparse attention, RAG integration, and production-grade deployment capabilities.

Tags

cuda · llm · gpu-optimization · transformer · quantization · inference · flashattention

Key entities

CUDA (technology, 0.99) · PyTorch (technology, 0.95) · FlashAttention (technology, 0.92) · HuggingFace (technology, 0.88) · Modal (technology, 0.9) · Safetensors (technology, 0.85) · RoPE (technology, 0.88) · RMSNorm (technology, 0.85) · SwiGLU (technology, 0.85) · BPE tokenizer (technology, 0.87) · LLaMA (technology, 0.9) · quantization (concept, 0.95) · sparse-attention (concept, 0.92) · RAG (concept, 0.9) · instruction-fine-tuning (concept, 0.88) · laptop-gpu (location, 0.85)

Classification

reference · language en · status final

Provenance

claude-haiku-4-5 via @stacklist/be@0.1.0, confidence 0.85, 17 Jul 2026