Loading…
2026 September 10-11 | Tokyo, Japan
View More Details & Registration

IMPORTANT NOTE: Timing of sessions and room locations are subject to change.
Type: Evaluation & Testing clear filter
Friday, September 11
 

13:30 JST

From Clicks To Context: Building an Open-Source Evaluation Pipeline for AI Agents - Inês Bolaños, PagerDuty
Friday September 11, 2026 13:30 - 13:55 JST
The AI industry has moved so fast that we are still evaluating probabilistic software using the same deterministic metrics we applied to traditional code. As a Product Analyst working on AI agents at PagerDuty, I saw a critical need for a new observability standard, one that moves beyond clicks to measure true reasoning and reliability. To address this, I’ve developed and open-sourced a specialized framework designed to help teams decide, with data, when to hire, train, or fire an AI agent. In this session, I will walk through the H.I.R.E. Framework methodology and share the technical architecture of an evaluation pipeline that turns qualitative conversational data into structured, actionable product insights. I will share the open-source repository containing these metric definitions and templates, providing resources for the community to move past agent-washing and toward building verifiable, trustworthy agentic systems.
Speakers
avatar for Inês Bolaños

Inês Bolaños

Senior Product Analyst, PagerDuty
Inês Bolaños focuses on the intersection of AI, product strategy and data reliability. With almost a decade of experience, she specializes in turning complex data into actionable product decisions. Combining a background in Communication with a Master’s in Big Data, Inês os helping... Read More →
Friday September 11, 2026 13:30 - 13:55 JST
Hall 1F

16:45 JST

Why Traditional AI Benchmarks Fail Voice Agents - Harshita Jain, Smallest AI
Friday September 11, 2026 16:45 - 17:10 JST
AI systems are increasingly evaluated using benchmarks designed for individual components: Word Error Rate (WER) for speech recognition, MOS for speech synthesis, latency for infrastructure, and task success for agents. Yet users experience none of these components in isolation when using voice agents, they experience conversations.

Voice agents expose the limitations of traditional evaluation more clearly than any other AI system. An agent can achieve state-of-the-art WER, MOS, and latency metrics while still delivering a frustrating user experience.

This talk explores why established metrics are becoming insufficient for real-time AI applications and introduces a framework for evaluating voice agents holistically. We will examine concepts such as semantic understanding versus transcription accuracy, latency distributions versus averages, interruption handling, recovery from errors, and conversation-level success metrics.

Through real-world examples from production voice systems, attendees will learn how to move beyond component benchmarks and begin measuring what ultimately matters: whether an AI system successfully helps a user achieve their goal.
Speakers
avatar for Harshita Jain

Harshita Jain

Developer Relations Engineer, Smallest AI
I’m a Developer Relations Engineer at Smallest AI, bridging real-time Voice AI and developer communities with focus on enabling developers to build with open-source orchestrators. Before this, I was a Software Engineer at Mobile Premier League where I built real-time networking... Read More →
Friday September 11, 2026 16:45 - 17:10 JST
Hall C
 
Share Modal

Share this link via

Or copy link

Filter sessions
Apply filters to sessions.