Anthropic's AI Risk Evaluators: Do They Have Too Little Power?
Proposed AI risk evaluators may not have enough power to stop increasingly capable models, according to experts. Anthropic CEO Dario Amodei has proposed embedding third-party safety evaluators inside frontier AI companies on an ongoing basis as part of a plan to slow the advance of large language models.
The proposal is similar to embedded bank supervisors in the banking industry, but experts say it's not accurate without comparable enforcement power or the legal authority to prevent a model from being trained or released. Amodei wrote that evaluators are needed to provide 'a neutral third party who can actually see the details,' but acknowledged that Anthropic still determines what it includes and omits from its existing public disclosures.
Anthropic's proposed evaluators would have access to frontier AI systems, but limited formal authority over the companies developing them. Neither Anthropic's proposal nor OpenAI's existing third-party evaluation framework gives outside evaluators independent authority to halt the development or deployment of a model.