{"version":"1.0","type":"card","id":"f19e38c4-896f-4094-ae19-08d4041ee6d1","url":"https://stacklist.com/card/f19e38c4-896f-4094-ae19-08d4041ee6d1","title":"The Three Futures of Voice AI (Substack)","source_url":"https://fdaudens.substack.com/p/the-three-futures-of-voice-ai","note":"Thirteen voice AI insiders from the Cerebral Valley Voice Summit cannot agree on where voice is headed, and that is the story. Three competing futures for voice: the platform shift, the plumbing layer, or the occasional feature.","image":{"url":"https://ucarecdn.com/7c6829da-bc6e-4aa3-9820-45255f3c56b6/","alt":"The Three Futures of Voice AI (Substack)","width":1456,"height":1078},"stack":{"id":"8521690e-e828-43bf-a05b-91287c60612b","title":"The Future of Voice","url":"https://stacklist.com/c/technology/stack/8521690e-e828-43bf-a05b-91287c60612b"},"created_at":"2026-07-20T17:38:57.438Z","updated_at":null,"aco":{"summary":"Voice AI insiders disagree on fundamental questions about the future of voice technology, including whether AI should sound human or distinctly artificial, and whether cascade or end-to-end architectures are superior. The consensus centers on latency as the critical technical barrier to overcome for natural human-AI interaction.","tags":["voice-ai","speech-recognition","latency","agentic-future","real-time-processing","ai-interaction"],"key_entities":[{"name":"Florent Daudens","type":"person","confidence":0.95},{"name":"Russ d'Sa","type":"person","confidence":0.9},{"name":"Brandon Yang","type":"person","confidence":0.9},{"name":"Justin Uberti","type":"person","confidence":0.9},{"name":"Scott Stephenson","type":"person","confidence":0.9},{"name":"Dylan Fox","type":"person","confidence":0.9},{"name":"OpenAI","type":"organization","confidence":0.95},{"name":"Deepgram","type":"organization","confidence":0.95},{"name":"AssemblyAI","type":"organization","confidence":0.95},{"name":"Cartesia","type":"organization","confidence":0.9},{"name":"LiveKit","type":"organization","confidence":0.9},{"name":"Sierra","type":"organization","confidence":0.85},{"name":"MiniMax","type":"organization","confidence":0.85},{"name":"Wispr Flow","type":"organization","confidence":0.85},{"name":"GPT-Realtime-2","type":"technology","confidence":0.95},{"name":"GPT-5","type":"technology","confidence":0.9},{"name":"latency","type":"concept","confidence":0.95},{"name":"uncanny valley","type":"concept","confidence":0.9},{"name":"auditory Turing test","type":"concept","confidence":0.9},{"name":"Cerebral Valley Voice Summit","type":"event","confidence":0.9},{"name":"Cerebral Valley","type":"location","confidence":0.85}],"classification":"analysis","language":"en","confidence":0.85,"provenance":{"model":"claude-haiku-4-5","tool":"@stacklist/be@0.1.0","confidence":0.85,"timestamp":"2026-07-20T17:39:13.315Z"},"token_counts":{"approximate":2379,"cl100k":2108},"content_hash":"sha256:7858589253923b68b97d1259beba5f82a5bbe6d6f0bb8cf84dd280d7495828b0","acp_version":"0.2","body_available":true,"body_tokens":2379,"visibility":"public","agent_accessible":true,"status":"final"},"_links":{"self":"/api/public/card/f19e38c4-896f-4094-ae19-08d4041ee6d1.json","html":"https://stacklist.com/card/f19e38c4-896f-4094-ae19-08d4041ee6d1","md":"/api/public/card/f19e38c4-896f-4094-ae19-08d4041ee6d1.md","stack_json":"/api/public/stack/8521690e-e828-43bf-a05b-91287c60612b.json"}}