Loading…
2026 September 10-11 | Tokyo, Japan
View More Details & Registration

IMPORTANT NOTE: Timing of sessions and room locations are subject to change.
Friday September 11, 2026 15:35 - 16:00 JST
Teams are moving coding agents onto self-hosted open-weight models to control inference cost, and real sessions routinely span tens of turns and hundreds of thousands of tokens. But the only common signal for judging those models is the static accuracy benchmark, which says little about how a model behaves across a real, multi-turn session.

This session closes that gap with a trace-based evaluation method built on open technology: OpenTelemetry for capture and MLflow for analysis. Through a case study tracing real Claude Code sessions on a 120B-class open-weight model and a frontier model, it surfaces behaviors no leaderboard reports. For example, in the sessions we traced, the model's logged reasoning recorded a user's constraint, and the same turn violated it. An accuracy score sees only the wrong action; the trace shows the model had the rule in hand and did not follow it.

Attendees leave with a reusable observability architecture for their own agents, metrics that go beyond accuracy, and a clear view of what running a coding agent on open weights actually costs, in reliability and not only in dollars.
Speakers
avatar for Kota Tsuyuzaki

Kota Tsuyuzaki

Engineering Manager, NTT DOCOMO BUSINESS, Inc.
Kota joined Nippon Telegraph and Telephone Corporation in 2010 and has been a core developer of OpenStack Swift, the open-source on-premise cloud storage. He later moved into the AI/HPC area, working with the Lustre file system and Slurm Workload Manager. In 2023 he joined NTT DOCOMO... Read More →
Friday September 11, 2026 15:35 - 16:00 JST
Hall B

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link