Meta finds two AI coding agents catch bugs better than one
New research from Meta highlights the significant benefits of using two AI coding agents for bug detection. The study found that pairing AI agents to review each other’s work substantially improves their ability to catch bugs compared to giving a single agent a larger budget.
Meta’s findings build on data from its internal review tool, RADAR, which has analyzed over 535,000 code changes. Of these, more than 331,000 were merged into the codebase. RADAR has also reduced the median review time by 35%. The tool uses risk calibration to adjust its caution based on the potential danger of code changes, resulting in a lower revert rate and a drastic reduction in production incidents compared to manual reviews.
Meta’s Engineering Agent was tested over a three-month period to repair test failures. Of the fixes it generated, 80% underwent human review, with about 25.5% of those reviewed fixes accepted into production. This shift suggests that engineers are increasingly focusing on evaluating machine-written fixes rather than writing them from scratch.
Additional testing frameworks, such as Meta’s Just-in-Time testing, combine large language models, program analysis, and mutation testing. Mutation testing intentionally introduces small faults to check if tests can detect them. Across over 22,000 generated tests, this approach led to a fourfold increase in bug detection, with improvements reaching up to twentyfold for significant failures.