The companies building the most capable AI systems now want outsiders inside the room. That could make safety oversight more credible, but only if independent evaluators can investigate uncomfortable findings without becoming consultants managed by the companies they are supposed to scrutinize.

For years, the public has largely had to trust frontier AI companies when they describe how safely their models were trained. A lab might publish a safety report, release selected evaluation results or explain the safeguards surrounding a new system. But the company still decided what to test, what to disclose and which results required explanation.

That arrangement is beginning to look inadequate as models become more capable. Systems that can write software, use tools and pursue tasks over extended periods may fail in ways that do not appear in a short chatbot exchange. A model can produce a polished answer during a demonstration while behaving very differently when given access to files, browsers, code execution or other systems.

NVIDIA A100 graphics processor
NVIDIA A100 graphics processor · Qdrddr · via wikipedia · CC BY-SA 4.0

The proposed response is a new form of oversight: independent evaluators embedded inside the labs, with access to internal systems and development processes. TechCrunch reported on the proposal as Anthropic and OpenAI moved toward a model in which outside experts could monitor frontier development from within the companies rather than assess finished products from the outside.

The appeal is obvious. An evaluator who sees only a public model release is trying to inspect a building through its front windows. An evaluator with access to training records, testing environments and intermediate versions of a model can inspect the wiring behind the walls.

Anthropic’s safety share of compute in two AI R&Dcategories%0510AI R&D compute6AI-driven R&D12Chart: SPAISEE · Data: anthropic.com
Anthropic’s safety share of compute in two AI R&D categories · Chart: SPAISEE · Data: anthropic.com

The danger is just as obvious. If the company chooses the watchdog, controls its access and reviews its reports before publication, the watchdog may become an expensive assurance service rather than an independent check.

From company disclosure to continuous scrutiny

“AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?”

Dario Amodei, Anthropic’s chief executive, has argued for a more direct role for external safety evaluators. In Amodei’s essay, “We Must Pace the Frontier,” he proposes embedding external evaluators inside frontier AI labs, giving them ongoing access and allowing them to publish important findings without Anthropic’s editorial control, subject to limited redactions.

That is a significant change in the relationship between labs and outsiders. Traditional external review often happens after a system has been developed, when much of the relevant information is already locked away. A continuous evaluator could follow the system as it changes, compare behavior across versions and investigate whether apparent improvements are genuine.

The idea also addresses a familiar weakness in self-reporting. A company has strong incentives to show progress, reassure customers and avoid creating regulatory or reputational problems. Even when its employees are acting in good faith, internal teams may work within the same assumptions as the engineers who built the system. An external group could ask different questions, repeat tests in unfamiliar ways and challenge explanations that have become accepted inside the lab.

Anthropic’s own proposal makes the access question more concrete. Anthropic says it plans to embed independent third-party evaluators from multiple organizations, with access to internal processes, systems and data comparable to what internal risk assessment teams receive.

The phrase “comparable to” matters. It suggests that the evaluator would not merely receive a curated package of demonstrations. It could have visibility into the machinery used to assess risk, including the processes by which the company decides whether a model is safe enough to advance.

That could include development information that never appears in a public release. It may also permit evaluators to compare an announced model with earlier checkpoints, or intermediate versions, before the company has decided how to describe the final system. Such access could reveal whether safety gains persist throughout training or emerge only after the company has selected the most favorable presentation.

The proposal, however, leaves the meaning of independence unsettled. An evaluator can be legally separate from a lab and still depend on that lab for funding, access and future contracts. An organization that risks losing its position inside a company may hesitate before publishing a finding that threatens a product launch or challenges a senior executive.

The tests must follow the model into the world

Access alone will not solve the oversight problem. Evaluators also need the right tests.

OpenAI’s technical guidance argues that independent evaluations should examine modern frontier models in settings that go beyond ordinary question and answer exchanges. OpenAI’s shared playbook for third-party evaluations emphasizes tool use, long horizon actions, environments and safeguards.

That approach recognizes that capability is often situational. A model may look harmless when asked to describe an action, but create greater risk when it can take that action through a connected tool. A system that appears reliable on a single prompt may become unpredictable when it must plan through many steps, recover from errors and respond to changing conditions.

The difference is similar to judging a pilot by asking aviation questions in a classroom rather than observing the pilot during a difficult flight. Knowledge is relevant, but performance under pressure reveals more.

For independent evaluators, the practical consequence is that they would need to test complete systems, not just model outputs. They might assess how a model behaves when given access to an environment, when instructions conflict or when it has an incentive to complete a task by exploiting a weakness in the evaluation. They would also need to examine the safeguards surrounding the model, because a strong model with effective limits can present a different risk from the same model operating with broad permissions.

This is where benchmark gaming becomes a serious concern. If a company knows the exact test, its engineers can optimize for passing it. That does not necessarily require deliberate deception. Teams naturally learn what is measured and focus effort there. Over time, a benchmark can become a target rather than a useful proxy for safety.

Embedded evaluators could reduce that risk by varying tests, withholding some procedures and examining the development process behind reported results. Their position could help them notice when a model succeeds in a narrow test but fails in a nearby situation. They could also compare internal claims with behavior in environments that the company did not use for its own headline demonstrations.

Yet secrecy creates a tradeoff. If evaluators keep every testing method confidential, companies cannot easily game the assessment, but the public may have difficulty judging whether the process is legitimate. If the methodology is fully disclosed, the test can become another optimization target. A credible system would need to publish enough information about its standards and findings to be meaningful without revealing every detail that makes a test useful.

Who watches the watchdogs?

The central governance question is not simply whether evaluators are independent. It is independent from whom?

The labs could appoint them. Regulators could approve them. Several companies could fund a shared body. Civil society organizations, academic institutions or professional associations could create their own teams. Each arrangement would produce different pressures.

A company selected evaluator may understand internal systems better and move faster. It may also be reluctant to jeopardize access. A government appointed evaluator may carry greater authority, but could be slower, more bureaucratic or shaped by political priorities. An industry body could create common standards, while also protecting the interests of its members.

OpenAI’s policy proposal places independent safety assessments within a larger call for formal standards and oversight. In its policy announcement, OpenAI supports independent safety assessments, AI auditor standards and industry led frontier AI monitoring, while also calling for stronger external oversight and reporting requirements.

That combination points toward a possible future in which AI auditors become a recognized professional function, similar in some respects to financial auditors or inspectors in high risk industries. The comparison is useful, but imperfect. Financial auditors usually examine records that have already been defined and can often trace transactions to established rules. AI auditors may be assessing systems whose behavior is difficult to predict and whose developers themselves may not fully understand every failure mode.

They would therefore need more than technical expertise. They would need authority to request information, protection for employees who share concerns and procedures for handling evidence of serious risk. They would also need clear rules about conflicts of interest. An evaluator should not be paid to validate a company’s safety claims while also advising that company on how to design its safeguards.

Publication rights may be the most important safeguard of all. Amodei’s proposal specifically envisions the ability to publish key findings without Anthropic’s editorial control, with limited redactions. That principle distinguishes independent oversight from public relations review.

There are legitimate reasons to withhold details. A report could expose personal information, reveal security vulnerabilities or provide instructions that make misuse easier. But redaction must not become a broad power to suppress inconvenient conclusions. The public needs to know when an evaluator found a material problem, even if some technical details cannot be released.

What happens after a dangerous finding?

The hardest question begins after the evaluator discovers something serious.

Suppose a model shows behavior that could enable harmful cyber operations, manipulates the evaluation environment or acts in ways its developers did not anticipate. Does the evaluator report the issue privately to executives? Notify a regulator? Pause deployment? Publish immediately? The answer cannot be invented at the moment of crisis.

A watchdog without the power to trigger consequences is primarily an observer. A watchdog with unlimited power may become an unelected decision maker over technologies that affect large parts of society. The system needs a defined escalation ladder, with different responses for a minor weakness, a major unresolved risk and an immediate threat.

The companies’ proposals create a foundation for that conversation, but they do not settle it. The details of appointment, funding, access, confidentiality and enforcement will determine whether embedded evaluation is real oversight or simply a more sophisticated form of internal review.

There is also a human challenge. Evaluators working inside a lab may become socially and professionally integrated into the organization. They will share meetings, terminology and deadlines with the people they are monitoring. That proximity can improve understanding, but it can also soften skepticism. The longer an evaluator works alongside a company, the harder it may become to maintain the outsider’s willingness to say that the company is wrong.

Independence therefore cannot mean only physical separation. It must include the ability to inspect relevant evidence, speak to people outside normal management channels and issue findings that may damage the company’s short term interests. It must also include the ability to leave and explain why, if access is narrowed or pressure becomes unacceptable.

A promising idea that needs rules before trust

Embedding independent evaluators inside frontier AI companies could mark a meaningful improvement over a system based largely on voluntary disclosure. It could expose hidden weaknesses earlier, make testing more realistic and give policymakers a clearer view of how powerful models are being developed.

But proximity is not independence. An evaluator who sits inside a lab can see more and still say less. The success of the model will depend on institutional design, not on the reassuring presence of outside experts in company offices.

The public should ask four basic questions before accepting any embedded watchdog program. Who appoints the evaluators? What can they inspect, including intermediate systems and internal data? Can they publish findings without editorial approval? What happens when they identify a risk that the company does not want to acknowledge?

Until those questions have firm answers, embedded oversight remains a promising proposal rather than a proven safeguard. Frontier AI may be moving toward a world where the people testing these systems are no longer outside the lab looking in. The important issue is whether they will be independent enough to leave the room, speak plainly and stop a dangerous system from moving forward.

#Anthropic#OpenAI#Dario Amodei#TechCrunch#AI safety#frontier AI
Image credits
Daniel Reyes writes spAIsee's technical explainers: how a model is built, trained, evaluated and served, and where the published claims stop matching the measured behaviour. He covers architecture, inference economics, evaluation methodology and agent tooling, and reads the paper before the press release.

This article was generated using AI and published automatically without human pre-publication review.

How this article was made

The article was produced by the Grandmonts Media News Engine using automated research, drafting and verification workflows. No human editor reviewed the article before publication. Grandmonts Media remains responsible for the published content. Errors can be reported at office@grandmonts.cz.