
044: Speech Recognition and Spoken Language Understanding (SLU)
About this episode
In our second technical dive in the the anatomy of a voice assistant, Kylie hosts Shawn Wen, co-founder and CTO of Poly AI, to analyze the complexities of building voice assistants. He highlights the challenges faced in speech recognition, such as dealing with errors, latency, and user experience. Shawn also discusses the inefficiencies of converting chatbots to voice assistants and the nuances that must be managed, including managing speech recognition models, right-sizing latency, agile dialogue design, and dealing with telephony filters. He emphasizes that optimizing these systems isn't just a technical problem but a multifaceted user experience challenge.
Get every episode summarized
Each time Deep Learning with PolyAI publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Deep Learning with PolyAI

What happens if your AI agents get 1% better every day?
Deep Learning with PolyAI

Who's coordinating your army of AI agents?
Deep Learning with PolyAI

Why should CX leaders care about MCP?
Deep Learning with PolyAI

Can AI really hear a call the way a person does?
Deep Learning with PolyAI