OpenAI Unveils New Framework for External Review of AI Models
OpenAI has announced a new framework for vetting its AI models before they are released to the public. The company is allowing independent evaluators to review its systems during earlier stages of development, with a focus on safety and potential risks.
The expanded evaluation program targets four priority areas: assessing OpenAI's 'safety cases,' probing the resilience of critical safeguards against adversarial threats, tying assessments directly to the company's Preparedness Framework, and investigating misalignment incidents.
The initiative builds on a commitment made by CEO Sam Altman, who signaled that evaluators would be embedded more deeply within OpenAI's operational structure. The new framework emphasizes scientific rigor, evaluator independence, security protocols, and responsible methods for publishing findings.