Anthropic and OpenAI need truly independent safety evaluators, experts say in public letter

Direct Source Verification: This story is aggregated from CNBC (cnbc.com). Full reporting rights and copyright belong to the primary publisher.
A coalition of over 100 AI experts are urging independence and transparency from Anthropic, OpenAI and other foundation model labs to conduct evaluations.

Over 100 artificial intelligence experts and evaluators are banding together to warn they won't have the necessary resources and protections to test the safety of AI technology, which is facing heightened scrutiny due to concerns from insiders about the potential dangers of frontier models.

"We're just trying to really demonstrate a shared common ground on basic principles and ensure that independent oversight can be a meaningful tool for managing AI risk broadly," said Conrad Stosz, chair of the AI Evaluator Forum consortium that organized the letter, in an interview.

The group published a public letter on the matter on Friday, and shared it exclusively with CNBC. The signatories include AI luminaries like Geoffrey Hinton and members of organizations such as Johns Hopkins University, Stanford University and the nonprofit evaluator METR. They want to compel foundation model providers to ensure that third-party AI evaluators are allowed the necessary "scientific objectivity, transparency, independence, and robust protections" to do their jobs effectively and credibly, the letter said.

Stosz said it's part of an effort to hold the foundation model companies accountable to their recent pledges to support more thorough third-party AI safety testing.

The niche community of evaluators has been catapulted into the limelight since Anthropic CEO Dario Amodei floated the idea over the weekend of providing some of them "employee-like access" to inspect and audit bleeding-edge foundation models and their development processes. While some industry leaders have called on the government to regulate AI development to ensure it's not spinning out of control, President Donald Trump and his former AI czar, David Sacks, have adamantly opposed such efforts.

Stosz said that the coalition doesn't "advocate for one particular way" to ensure that AI models are developed safely, but wants to ensure that "basic principles" and "greater standardization" are at least established for evaluators and others who work independently of the major labs.

Amodei's proposal, Stosz said, appears to involve providing significantly more access than evaluators have previously enjoyed. Such a scenario, Stosz said, could involve foundation model makers giving third-party evaluators access to company computers, allowing them to talk to employees candidly and letting them "see sensitive internal data and unreleased systems."

"That type of access would give us much greater confidence and certainty about the actual risk, particularly for systems that they're using internally and not releasing," Stosz said. He cited the unreleased OpenAI model used in the Hugging Face attack.

OpenAI CEO Sam Altman, SpaceX's Elon Musk and Microsoft CEO Satya Nadella have publicly supported Amodei's proposal, but they've yet to address the logistical issues with such an undertaking, such as which AI evaluators will be selected and how deeply they would get to inspect closely guarded technologies.

The signatories want the work of evaluators to be conducted independent from the businesses, with more transparency about the technologies, and "to be shielded from retaliation from the companies they embed with," the letter said.

Vinh Nguyen, a Council on Foreign Relations senior fellow for AI and former chief AI officer of the National Security Agency, said independent evaluators are needed to help unearth crucial information that could help mitigate potential security failures and economic calamities.

"When a few powerful labs control capabilities that can endanger the cybersecurity, critical infrastructure, and the systems our national security and economy run on, the government and the public cannot be dependent on those labs' own account of what's secure and safe," Nguyen, who signed the letter, said in a statement.

Stosz said third-party evaluators are not intended to "be a replacement for any internal efforts to evaluate, let alone mitigate issues that that developers find." He acknowledged that it's possible the foundation model companies ignore the public letter and the call to action, but said their credibility is at stake.

"There's a very small number of groups that are actually sufficiently technically credible and have the scale and the ability" to perform the kind of work, he said.

Minimum Conditions for Embedding EvaluatorsWe, the undersigned, are encouraged to see frontier AI companies call for embedding third-party organizations to evaluate rapidly escalating AI capabilities and risks. We believe that all frontier AI companies should embed evaluators to independently assess AI risks, including evaluating the systems themselves and any significant incidents of real-world harm, as well as the companies' training, deployment, oversight, operational, and safeguard practices.To be credible, embedded third-party evaluations must have scientific objectivity, transparency, independence, and robust protections against interference from the evaluated companies, including at least:

WATCH: Presentation of AI technology to public "could not be worse."

Original Source
https://www.cnbc.com/2026/09/18/ai-safety-evaluators-anthropic-openai-models-security.html
Visit CNBC β†—
SHARE STORY:
𝕏 f in

Related Coverage in Business