Loading…
2026 September 10-11 | Tokyo, Japan
View More Details & Registration

IMPORTANT NOTE: Timing of sessions and room locations are subject to change.
Type: Evals & Testing clear filter
Friday, September 11
 

10:20 JST

Building Local Regression Testing and Evals for ADK Agents - Thu Ya Kyaw, Google
Friday September 11, 2026 10:20 - 10:45 JST
The hardest part of agentic engineering isn't writing the code; It’s proving that a prompt tweak or a new tool schema didn't completely break your agent's routing logic. Traditional unit testing falls short when dealing with non-deterministic agent trajectories.

This session looks at how to build an automated local evaluation and regression testing pipeline for ADK workflows. We will walk through how to programmatically mock tool responses, simulate user edge cases, and run parallel assertions against agent trajectories using open-source evaluation frameworks. Attendees will learn how to catch infinite loops, detect tool-calling degradation, and establish a baseline scoring rubric for agent accuracy before code ever hits a production branch.

___________________________
Presentation Language: English

Captioning will be available for attendees in 50+ languages through Wordly. See instructions in each room to utilize captioning.

Speakers
avatar for Thu Ya Kyaw

Thu Ya Kyaw

Senior Developer Relations Engineer, Google
Thu Ya Kyaw is a Senior Developer Relations Engineer for Google Cloud. At Google, he helps to make learning, developing, deploying, and scaling applications on Google Cloud a delightful experience for everyone. He is passionate about using AI to solve real-world problems, and he is... Read More →
Friday September 11, 2026 10:20 - 10:45 JST
Hall C
  Evals & Testing

13:40 JST

From Clicks To Context: Building an Open-Source Evaluation Pipeline for AI Agents - Inês Bolaños, PagerDuty
Friday September 11, 2026 13:40 - 14:05 JST
The AI industry has moved so fast that we are still evaluating probabilistic software using the same deterministic metrics we applied to traditional code. As a Product Analyst working on AI agents at PagerDuty, I saw a critical need for a new observability standard, one that moves beyond clicks to measure true reasoning and reliability. To address this, I’ve developed and open-sourced a specialized framework designed to help teams decide, with data, when to hire, train, or fire an AI agent. In this session, I will walk through the H.I.R.E. Framework methodology and share the technical architecture of an evaluation pipeline that turns qualitative conversational data into structured, actionable product insights. I will share the open-source repository containing these metric definitions and templates, providing resources for the community to move past agent-washing and toward building verifiable, trustworthy agentic systems.

___________________________
Presentation Language: English

Captioning will be available for attendees in 50+ languages through Wordly. See instructions in each room to utilize captioning.
Speakers
avatar for Inês Bolaños

Inês Bolaños

Senior Product Analyst, PagerDuty
Inês Bolaños focuses on the intersection of AI, product strategy and data reliability. With almost a decade of experience, she specializes in turning complex data into actionable product decisions. Combining a background in Communication with a Master’s in Big Data, Inês is helping... Read More →
Friday September 11, 2026 13:40 - 14:05 JST
Hall 1F
  Evals & Testing

15:35 JST

Trace-Based Evaluation of Open-Weight Coding Agents: Measuring Real Agent Behavior - Kota Tsuyuzaki, NTT DOCOMO BUSINESS, Inc.
Friday September 11, 2026 15:35 - 16:00 JST
Teams are moving coding agents onto self-hosted open-weight models to control inference cost, and real sessions routinely span tens of turns and hundreds of thousands of tokens. But the only common signal for judging those models is the static accuracy benchmark, which says little about how a model behaves across a real, multi-turn session.

This session closes that gap with a trace-based evaluation method built on open technology: OpenTelemetry for capture and MLflow for analysis. Through a case study tracing real Claude Code sessions on a 120B-class open-weight model and a frontier model, it surfaces behaviors no leaderboard reports. For example, in the sessions we traced, the model's logged reasoning recorded a user's constraint, and the same turn violated it. An accuracy score sees only the wrong action; the trace shows the model had the rule in hand and did not follow it.

Attendees leave with a reusable observability architecture for their own agents, metrics that go beyond accuracy, and a clear view of what running a coding agent on open weights actually costs, in reliability and not only in dollars.

___________________________
Presentation Language: English

Captioning will be available for attendees in 50+ languages through Wordly. See instructions in each room to utilize captioning.
Speakers
avatar for Kota Tsuyuzaki

Kota Tsuyuzaki

Engineering Manager, NTT DOCOMO BUSINESS, Inc.
Kota joined Nippon Telegraph and Telephone Corporation in 2010 and has been a core developer of OpenStack Swift, the open-source on-premise cloud storage. He later moved into the AI/HPC area, working with the Lustre file system and Slurm Workload Manager. In 2023 he joined NTT DOCOMO... Read More →
Friday September 11, 2026 15:35 - 16:00 JST
Hall B
  Evals & Testing
 
Share Modal

Share this link via

Or copy link

Filter sessions
Apply filters to sessions.