Loading…
2026 September 10-11 | Tokyo, Japan
View More Details & Registration

IMPORTANT NOTE: Timing of sessions and room locations are subject to change.
Friday September 11, 2026 16:45 - 17:10 JST
AI systems are increasingly evaluated using benchmarks designed for individual components: Word Error Rate (WER) for speech recognition, MOS for speech synthesis, latency for infrastructure, and task success for agents. Yet users experience none of these components in isolation when using voice agents, they experience conversations.

Voice agents expose the limitations of traditional evaluation more clearly than any other AI system. An agent can achieve state-of-the-art WER, MOS, and latency metrics while still delivering a frustrating user experience.

This talk explores why established metrics are becoming insufficient for real-time AI applications and introduces a framework for evaluating voice agents holistically. We will examine concepts such as semantic understanding versus transcription accuracy, latency distributions versus averages, interruption handling, recovery from errors, and conversation-level success metrics.

Through real-world examples from production voice systems, attendees will learn how to move beyond component benchmarks and begin measuring what ultimately matters: whether an AI system successfully helps a user achieve their goal.
Speakers
avatar for Harshita Jain

Harshita Jain

Developer Relations Engineer, Smallest AI
I’m a Developer Relations Engineer at Smallest AI, bridging real-time Voice AI and developer communities with focus on enabling developers to build with open-source orchestrators. Before this, I was a Software Engineer at Mobile Premier League where I built real-time networking... Read More →
Friday September 11, 2026 16:45 - 17:10 JST
Hall C

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link