OpenAI and Anthropic came close to a legally binding agreement that would have required each company to stress test the other’s artificial intelligence models, according to an exclusive report from The Information. The discussions reportedly began before a series of security incidents intensified concerns within OpenAI, adding a new layer of urgency to an already competitive relationship.

The proposed arrangement would have created an unusual form of cooperation between two companies competing to build some of the world’s most capable AI systems. Rather than relying solely on their own safety teams, OpenAI and Anthropic would have given each other a formal role in looking for vulnerabilities, dangerous behaviors and failures that might remain hidden during internal reviews.

The Information did not report that the agreement was signed. It also did not make clear whether the talks were abandoned, remain unresolved or could be revived. Details about the proposed testing methods, the information the companies would have shared and the legal obligations attached to the arrangement were not disclosed.

1515 Third Street
1515 Third Street · Coolcaesar · via wikipedia · CC BY 4.0

Even so, the discussions point to a growing recognition inside the AI industry that internal testing has limits. Developers routinely evaluate their models for inaccurate answers, harmful instructions, privacy problems and attempts to bypass safety controls. But a company testing its own product may not always identify the weaknesses that an outside team would find.

Rivalry creates a testing problem

OpenAI and Anthropic are commercial rivals, but they also face many of the same technical and social risks. Their models are increasingly used to write software, analyze confidential documents, operate tools and interact with computer systems. As these systems become more capable, a failure can involve more than an incorrect response. It could expose sensitive information, misuse access permissions or help a user carry out harmful activity.

Cross testing could provide a form of independent scrutiny without requiring either company to surrender control of its models. Each lab could search for weaknesses in the other’s systems, report its findings and potentially help establish whether the problems were fixed. In principle, that would resemble the way security researchers test software made by companies that are otherwise competitors.

The arrangement would also have allowed the companies to compare their definitions of risk. One lab might consider a model’s ability to manipulate a user a serious warning sign, while another might focus more heavily on privacy or cybersecurity. Shared testing could reveal those differences and make it easier for customers, researchers and regulators to understand what safety claims actually mean.

Security incidents raise the stakes

The talks reportedly took place before a series of AI related security incidents increased concern inside OpenAI. The post did not specify the incidents, their impact or whether they were directly connected to the proposed agreement. Still, the timing is significant.

Security problems often expose weaknesses that standard product testing misses. A model may appear safe in a controlled evaluation but behave differently when connected to external tools, supplied with carefully designed prompts or placed in an unfamiliar environment. Attackers also have an incentive to search for combinations of weaknesses that internal teams may not anticipate.

A legally binding agreement could have given the testing process more structure. It might have specified how findings would be kept confidential, how quickly serious vulnerabilities must be disclosed and what would happen if one company believed the other had failed to respond. It could also have addressed liability, access to sensitive systems and disputes over whether a reported problem was genuine.

A test of voluntary oversight

The proposal reflects a broader question facing the AI industry: can companies voluntarily cooperate on safety while competing for customers, talent and market share?

Joint testing could strengthen public confidence, particularly for businesses considering whether to place private data or important operations in AI systems. But cooperation would be difficult to sustain if either company feared that sharing information could weaken its commercial position or reveal an embarrassing flaw.

For regulators, the talks may offer a useful early example of both the promise and limits of industry self governance. A cross testing pact would not replace independent oversight, but it could provide practical evidence that rival laboratories are willing to challenge one another’s claims.

The fact that the companies nearly reached such an agreement is therefore important, even without confirmation that it was completed. It suggests that the risks of advanced AI are becoming too complex for any one company to assess alone. It also shows how difficult it may be to turn that recognition into durable cooperation when the same companies are racing to lead the market.

#OpenAI#Anthropic#The Information#AI safety#AI security#model testing
Image credits
Daniel Reyes writes spAIsee's technical explainers: how a model is built, trained, evaluated and served, and where the published claims stop matching the measured behaviour. He covers architecture, inference economics, evaluation methodology and agent tooling, and reads the paper before the press release.

This article was written with the assistance of an AI system and published automatically.