AI Grading Its Own Homework: Double-Blind Pilot Yields Surprising Results
The world's first double-blind AI evaluation pilot was recently conducted at a massive scale. The AAAI-26 conference, one of the premier venues for AI research, used artificial intelligence to grade its own homework in the form of paper submissions. A total of 22,977 main-track papers were processed using an AI-powered peer review system.
The AI reviews were produced by a multi-stage LLM-based system that could process papers in under one day. The AI-generated reviews identified themselves as such and were slotted in alongside at least two human reviews for each paper. This double-blind framework ensured that both authors and reviewers remained anonymous.
Surveys conducted after the pilot found that both authors and program committee members preferred the AI-generated reviews over their human counterparts. They rated them higher on technical accuracy and quality of research suggestions. The AI-powered peer review process was seen as a supplement to, rather than a replacement for, human reviewers.