{"version":"1.0","type":"card","id":"faffd203-636d-465e-8ac4-f0c65b38ebcf","url":"https://stacklist.com/card/faffd203-636d-465e-8ac4-f0c65b38ebcf","title":"If you're still training in FP16, you're leaving half your GPU on the table.","source_url":"https://www.linkedin.com/posts/paoloperrone_if-youre-still-training-in-fp16-youre-share-7480370913650176000-CIDz/?utm_source=share&utm_medium=member_ios&rcm=ACoAAAI21ZsBNnZPaKuTab7nquKLCveUW7o-1DE","note":"This page discusses the advantages of using 8-bit floating point formats over FP16 for training neural networks. It highlights the benefits of E4M3 and E5M2 formats for different training phases, emphasizing memory savings, increased throughput, and maintaining accuracy.","image":{"url":"https://ucarecdn.com/edcbbbcc-4314-4a7f-bdea-ede67799f074/","alt":"If you're still training in FP16, you're leaving half your GPU on the table.","width":1280,"height":800},"stack":{"id":"4aae218c-38d7-4c05-b7ae-f2db0029a2e8","title":"Local AI & GPUs","url":"https://stacklist.com/stack/4aae218c-38d7-4c05-b7ae-f2db0029a2e8"},"created_at":"2026-07-07T23:26:05.677Z","updated_at":null,"aco":{"summary":"FP8 floating-point training offers 2x memory savings and throughput improvements over FP16 by using two specialized formats (E4M3 for weights and E5M2 for gradients) that NVIDIA's Transformer Engine switches between automatically. The post discusses the evolution from FP32 to FP8 and hints at emerging NVFP4 technology that further optimizes precision reduction while preserving network accuracy.","tags":["fp8-training","gpu-optimization","mixed-precision","transformer-engine","neural-networks","model-compression","deep-learning"],"key_entities":[{"name":"Paolo Perrone","type":"person","confidence":0.95},{"name":"Michał Piszczek","type":"person","confidence":0.95},{"name":"Curtis Burkhalter","type":"person","confidence":0.95},{"name":"NVIDIA","type":"organization","confidence":0.98},{"name":"Transformer Engine","type":"technology","confidence":0.98},{"name":"H100 Tensor Cores","type":"technology","confidence":0.95},{"name":"Hopper","type":"technology","confidence":0.92},{"name":"Blackwell","type":"technology","confidence":0.9},{"name":"NVFP4","type":"technology","confidence":0.93},{"name":"FP8 training","type":"concept","confidence":0.99},{"name":"mixed-precision training","type":"concept","confidence":0.98},{"name":"E4M3 format","type":"concept","confidence":0.95},{"name":"E5M2 format","type":"concept","confidence":0.95}],"classification":"analysis","language":"en","confidence":0.85,"provenance":{"model":"claude-haiku-4-5","tool":"@stacklist/be@0.1.0","confidence":0.85,"timestamp":"2026-07-07T23:26:11.860Z"},"token_counts":{"approximate":579,"cl100k":545},"content_hash":"sha256:2e2f9af8375a6ea480351eab34379348fe67b1a2751145e29401d3e6c401908b","acp_version":"0.2","body_available":true,"body_tokens":579,"visibility":"public","agent_accessible":true,"status":"final"},"_links":{"self":"/api/public/card/faffd203-636d-465e-8ac4-f0c65b38ebcf.json","html":"https://stacklist.com/card/faffd203-636d-465e-8ac4-f0c65b38ebcf","md":"/api/public/card/faffd203-636d-465e-8ac4-f0c65b38ebcf.md","stack_json":"/api/public/stack/4aae218c-38d7-4c05-b7ae-f2db0029a2e8.json"}}