Microsoft's AI Diagnostic System Crushes Human Physicians in Medical Case Challenge
Microsoft's AI diagnostic system has made headlines for its impressive accuracy in solving difficult medical cases. In June 2025, the company reported that its experimental diagnostic orchestrator, paired with OpenAI's o3 model, correctly solved as many as 85.5% of 304 exceptionally difficult medical cases adapted from the New England Journal of Medicine.
The results were released as a preprint and have since been revised to reflect new data. In November 2025, the maximum accuracy of the AI system was reported at 84.5%, while physicians averaged 36.1% on the same cases. The system's performance was tested on a controlled, text-based benchmark, not on real patients.
The benchmark, called SDBench, consisted of 304 consecutive clinicopathological conference cases published between 2017 and 2025. These cases were selected for their difficulty and rarity, making the exercise useful for testing diagnostic breadth but also limiting its relevance to everyday clinical practice.