Microsoft and Carnegie Mellon Unite to Define Standards for Evaluating AI Agents
The field of artificial intelligence (AI) is rapidly advancing, but there's a significant gap in our ability to evaluate and govern AI agents. To address this issue, Microsoft Research and Carnegie Mellon University's AI Measurement Science & Engineering Center (AIMSEC) are hosting a two-day workshop on October 22-23, 2026.
The workshop will bring together 120 key leaders from academia, industry, civil society, and government to define shared standards and measurement science for evaluating and governing AI agents. The program includes keynote presentations, lightning talks, panel discussions, and hands-on working groups addressing AI agent evaluation frameworks, tooling, and domain-specific challenges.
The workshop aims to create a foundation of accountability by developing rigorous, shared measurement science. This will enable the development of standards and best practices for evaluating and governing AI agents. The event also seeks to foster an ongoing cross-sector collaboration on agent evaluation, seeding working groups and partnerships that carry the work forward.