Users Take Center Stage in Microsoft AI Evaluation Overhaul
Microsoft AI researchers are working to improve AI evaluation methods by putting users back at the center of the process. The team, led by Principal UX Researchers Christopher Monier and Chuck Kwong, along with Wendy Wang, has developed a response-quality evaluation program that compares Copilot's answers against other AI products.
The researchers noticed a gap between what teams thought was a good response and what users actually preferred. They conducted qualitative conversations with real users to understand the underlying values that people truly care about. One of the dimensions they derived is accuracy, which is difficult for machines to evaluate but crucial for users.
Wendy Wang's team takes it a step further by turning data into 'golden' datasets and building shared vocabularies. They also use multi-turn evaluations, where people interact with the AI and evaluate the whole exchange rather than a single response.