A benchmark score has become one of the simplest ways to compare frontier models. Companies publish tables, investors cite rankings and customers use individual results as evidence that one system is more capable than another. Yet the value of those numbers depends on an assumption that is becoming harder to defend: the model has not already encountered the test.
That assumption is under pressure from the way modern models are trained. Public benchmarks can enter training data, appear in research papers, circulate in evaluation suites or be reproduced in countless online discussions. Even when a test has not been directly included in training, models can be optimized against similar tasks. The result is a familiar competitive problem. Labs may be improving their systems, but they may also be improving their ability to perform on known exams.
Google DeepMind is now testing a way to make that problem harder to ignore. In its announcement, Google DeepMind said it is piloting double blind AI evaluations that use cryptographically secured environments to assess a model against confidential tests. The pilot evaluates Gemini 2.5 Flash Lite while keeping the model weights and evaluation prompts from the other party.
The significance is not simply that DeepMind has introduced another benchmark. It is that the company is treating access to the exam as a security problem. If the questions are valuable enough to determine how a frontier model is judged, they must be protected like sensitive data.
The problem with a familiar exam
A public benchmark is useful because other researchers can inspect it, reproduce results and compare systems over time. Openness creates a common language for the industry. It also creates an attack surface.
Once questions are public, developers can test models against them during training or post training optimization. A benchmark can remain widely used while becoming less informative. High performance may reflect genuine generalization, but it may also reflect familiarity with the task format, the data distribution or the exact questions.
This does not make every public score meaningless. It changes what the score can support. A result on a well known test may show that a system performs strongly under those published conditions. It does not necessarily establish that the model will perform equally well on unfamiliar problems.
The industry has responded in several ways. Labs have created internal tests, added newly written evaluations and reported results on tasks that are not publicly available. Those approaches may produce better evidence, but they introduce a different weakness. Outsiders cannot easily inspect a private benchmark, verify how it was administered or determine whether the company selected favorable questions.
The central dilemma is therefore symmetrical. A benchmark must be private to resist contamination, but an evaluation must be observable enough to deserve trust. DeepMind’s pilot is designed around that tension.
A secure room for the exam
The technical report, “Double Blind Evals: Resolving the Dual Confidentiality Dilemma in AI Safety Auditing,” describes the evaluation framework. Its core design uses secure enclaves and cryptographic protections to allow an evaluation to take place without exposing confidential information to either side.
In practical terms, the model owner should not receive the private test prompts, while the party providing the test should not receive the model’s weights. That separation matters because both assets are commercially and strategically sensitive. Test questions can lose their value if they become public. Model weights can reveal the system itself or provide a route to unauthorized replication.
The enclave acts as a controlled room in which the evaluation can run. The model and the prompts are brought together for the test, but the parties outside the environment are restricted from inspecting the protected contents. The report describes cryptographic guarantees intended to make the process auditable without requiring either party to surrender its most sensitive material.
This is not the same as making a benchmark fully open. Instead, it tries to make the procedure verifiable while keeping the questions confidential. That distinction could become important as evaluations move beyond academic comparison and into safety auditing, procurement and regulation.
The pilot reportedly involved Gemini 2.5 Flash Lite and private benchmarks from MLCommons and Singapore’s AI Safety Institute. Those tests are significant in the context of the experiment because they are not simply another set of public questions placed inside a secure environment. They represent an attempt to evaluate a frontier model using assessments that remain controlled by their owners.
The approach also changes the role of the benchmark provider. Rather than handing questions to a company and accepting its report, the provider can participate in an evaluation where the questions remain protected. The model developer, in turn, can obtain an assessment without gaining access to the test set. Neither side must rely entirely on the other’s internal assurances.
A different reading of frontier claims
The timing matters because model companies increasingly publish several kinds of results at once. A launch may include performance on established public benchmarks, private internal tests and newly created evaluations. Those categories can all be useful, but they do not answer the same question.
A public score offers comparability. An internal score may offer realism or depth, but the audience must trust the company’s process. A new evaluation can reduce exposure to historical test data, although its novelty and difficulty may be harder for outsiders to judge.
OpenAI’s announcement for GPT 6 Astra presents results across public, internal and newly created evaluations, while also discussing measures intended to address exposure to historical benchmark data. That combination reflects the industry’s emerging compromise. Companies still need public tests to establish continuity with earlier systems, but they increasingly need private or newly constructed tests to argue that the results measure more than memorization.
OpenAI’s accompanying GPT 6 Astra system card documents the model’s capability, safety and evaluation methodology, including information relevant to interpreting its benchmark and cybersecurity claims. The existence of those multiple evaluation layers is itself revealing. No single test is expected to carry the full burden of proving what a frontier model can do.
The comparison with DeepMind’s pilot is not a simple contest between one company’s methodology and another’s. OpenAI’s approach, as described in the supplied materials, focuses on reporting results across different evaluation categories and on measures related to historical benchmark exposure. DeepMind’s pilot focuses more directly on the confidentiality of the test and the conditions under which an independent party can examine a model.
Both approaches recognize the same underlying risk. Benchmarking is no longer only a measurement exercise. It is also an information security exercise.
Who controls the exam?
Double blind testing may improve confidentiality, but it does not eliminate questions about authority.
The first is access. If private evaluations become the most trusted evidence, who is allowed to administer them? A small group of model companies, research organizations or government backed institutes could become the gatekeepers of the tests that shape public perception. That could reduce contamination while concentrating influence.
The second is reproducibility. Researchers value benchmarks partly because they can repeat an experiment. A confidential evaluation cannot be reproduced in the ordinary sense if the questions remain hidden. A secure protocol may show that the test was administered as claimed, but it cannot give every outside researcher the ability to run the same questions.
The third is selection. A cryptographically protected benchmark can prevent a model from seeing the questions, but encryption cannot determine whether the questions are representative, sufficiently difficult or relevant to real world use. It can protect an exam without proving that the exam is good.
That places more responsibility on benchmark designers. They will need to publish methodology, sampling rules, scoring procedures and governance details without revealing the questions themselves. They may also need to use multiple test providers so that no single institution controls the definition of progress.
There is a further commercial question. A model company may be willing to submit to a confidential evaluation when the result is favorable, but less willing when the process is controlled by a competitor or when the outcome could affect sales and regulation. Participation will depend on whether the evaluation is seen as neutral, technically secure and consequential enough to justify the risk.
From scoreboards to evidence
The next stage of AI benchmarking is therefore unlikely to be defined by one perfect replacement for public tests. More likely, the industry will build layers of evidence.
Public benchmarks will remain useful for historical comparisons and independent experimentation. Fresh internal tests will help labs evaluate capabilities that are difficult to measure with older datasets. Confidential double blind assessments may provide a stronger check against contamination, especially where the benchmark provider must protect sensitive safety or security tasks.
The most credible model reports may eventually combine all three. They could show how a system performs on known public tasks, how it performs on new private tasks and whether an external evaluator can verify the process without receiving the model weights or test prompts.
That would shift attention away from record scores alone. A result would be judged by its provenance: who wrote the questions, who administered the test, what protections were used, whether the model developer had access to the data and how much of the procedure can be independently checked.
This is a more complicated story than a leaderboard. It is also closer to what customers and policymakers actually need. A business choosing a model does not only want to know whether it topped an exam. It wants evidence that the result predicts performance on unfamiliar work. A safety authority does not only want a company’s best score. It wants confidence that the evaluation was not shaped around the system being tested.
DeepMind’s pilot does not settle those questions. It does, however, identify where the competition is heading. As public benchmarks become easier to study and optimize against, control over the testing environment may become a strategic asset. The frontier model race could increasingly be decided not by who can produce the highest number, but by who can produce the most credible number under conditions that competitors cannot quietly prepare for.
The benchmark wars are moving behind closed doors. The industry’s next challenge will be proving that the doors are secure, and that someone independent still has the key.
- Gciriani · CC BY-SA 4.0
This article was generated using AI and published automatically without human pre-publication review.
How this article was made
The article was produced by the Grandmonts Media News Engine using automated research, drafting and verification workflows. No human editor reviewed the article before publication. Grandmonts Media remains responsible for the published content. Errors can be reported at office@grandmonts.cz.