Code Review Bottleneck: Autonomous Agents Write More Code, Ship Less
Laurie Voss, Head of Developer Relations at Arize AI and co-founder of npm, has made a compelling argument that code review is not dying but being rebuilt as an engineered system. According to a study involving over 100,000 GitHub developers, teams using autonomous agents wrote 741% more code but shipped only 30% more software. This bottleneck is caused by human review.
Voss cited a two-decade-old Cisco study finding that reviewers stop effectively finding defects beyond 400 lines per sitting. At this pace, a single 10,000-line agent pull request would take three to four working days of genuine human review. Voss argued that the solution is not to review harder but to stop reviewing PRs directly and instead invest human judgment in building automated review harnesses.
One such example is OpenAI's internal product built with no manually written code, where agents wrote everything: roughly one million lines of code, about 1,500 merged pull requests, built by three engineers in five months. However, the Bun case illustrates the verification gap, as a port passed 99.8% of its test suite while containing three orders of magnitude more memory-unsafe assertions than a comparable human-written Rust codebase.