AI Models Still Can't Do Real Work: Mercor's Industrialized Measurement Loop
Brendan Foody, CEO of Mercor, argues that the next breakthrough in AI is teaching models to use real software on laptops, not just pass tests. In a recent presentation, Foody discussed how 'RL environments' have become the primary training ground for frontier AI agents. These environments are high-fidelity clones of applications like Salesforce and Microsoft 365, populated with real-world data and scored by human experts.
Mercor's network has contributed 2.5 million expert hours in the second quarter of 2026 alone, powering post-training runs that lifted a corporate-law benchmark from 4.7% to 26.6%. Foody framed this as the industrialization of a measurement loop, where experts design environments, models roll out synthetic trajectories, and human reviewers validate scores.
Foody also emphasized the importance of humans in the training loop, arguing that they are essential for measuring what is beyond the frontier of model capabilities. He noted that models struggle to identify their own mistakes, a phenomenon he called the 'homework problem.'