As public benchmarks lose influence, Vals is betting that private, job specific evaluations will become the standard for deciding whether artificial intelligence is ready for professional and high stakes use. Its $40 million funding round gives that argument new weight, while raising difficult questions about transparency, trust and who gets to define model quality.

From scores to work

For years, AI progress was communicated through public leaderboards. A model that scored higher on a mathematics test, a coding suite or a collection of general knowledge questions could claim a measurable advantage over its competitors. Those comparisons helped researchers track rapid improvements, but they are becoming less persuasive as models grow more capable and developers become better at optimizing for the tests themselves.

Vals is built around the idea that the next generation of evaluations must look less like school examinations and more like professional work. The San Francisco startup keeps its specific testing materials private and concentrates on tasks in fields including law, finance and software development. Its goal is not simply to determine whether a model can produce a correct answer. It is to assess whether the output is useful, reliable and comparable to work produced by a human in a real setting.

That distinction matters. A model may perform well on a public coding benchmark while struggling to maintain a large software project, interpret ambiguous requirements or avoid introducing security vulnerabilities. It may answer legal questions accurately in isolation but fail to identify a conflict, cite the wrong authority or communicate uncertainty to a client. In finance, a polished response can still be dangerous if it misses an assumption that changes the conclusion.

Private testing allows Vals to examine those broader qualities without giving model developers a fixed target to optimize against. It also allows the company to update its evaluations as systems improve. A public test can become less informative when its questions appear in training data or when companies design specifically for its scoring system.

A larger definition of risk

Vals says its evaluations cover more than capability. The company also examines negative outcomes and the risks that could emerge when models are deployed in the world. Its stated testing portfolio includes recursive self improvement, mental health, cybersecurity, biosecurity and the law of armed conflict, including questions involving the Geneva Convention.

The breadth of those subjects points to an important change in the AI market. Benchmarking is no longer only about ranking models by intelligence. It is increasingly about establishing whether they can be trusted in contexts where mistakes carry financial, legal, social or physical consequences.

That creates a more complicated measurement problem. A system that gives a wrong answer on a trivia test is easy to score. A system that gives dangerous advice, exposes sensitive information or helps a user carry out a harmful act requires a framework that accounts for context and severity. Evaluators must consider not only whether an answer is technically correct, but also how it might be used and whether the model appropriately refuses or limits its response.

Vals’ approach could therefore become relevant to companies deciding which systems to deploy, as well as to regulators examining AI used in high stakes environments. If a benchmark can show that a model performs well on realistic tasks while avoiding unacceptable behaviors, it may carry more commercial value than another public score on an academic dataset.

The business of independent judgment

The company’s financing reflects investor confidence in that opportunity. Vals, founded in 2024 by Rayan Krishnan, has raised $40 million in a Series A led by Andreessen Horowitz. Its earlier seed round was led by 8VC and Bloomberg Beta. The company says revenue has grown to eight times last year’s level, while its staff has increased from eight people to 25.

Vals has also launched a program to provide model evaluations to federal agencies. That move could expand the importance of its work beyond private procurement. Government agencies are among the most consequential buyers of AI systems, but they also face requirements around accountability, security and public confidence. A private evaluation provider could help them compare systems more quickly, particularly in specialized areas where internal testing resources are limited.

The opportunity comes with a potential conflict. Benchmark companies are paid by organizations that want to understand or promote the performance of their models. That does not necessarily compromise the results, but it makes methodology and independence central to the product. If customers cannot inspect the tests, they must trust the evaluator’s design, scoring and governance.

Vals argues that secrecy protects the integrity of its evaluations. The counterargument is that undisclosed tests are difficult to reproduce or challenge. Researchers cannot easily determine whether a result reflects genuine capability, a narrow task design or an evaluator’s assumptions about what good performance means.

What buyers will believe

This tension may define the next stage of AI benchmarking. Public tests offer openness, but they can become stale or vulnerable to training contamination. Private tests offer freshness and resistance to gaming, but their conclusions are harder for outsiders to verify.

The market may eventually adopt a layered approach. Public benchmarks could continue to provide common reference points for research, while private evaluations measure whether a model is ready for a specific company, profession or government use. Buyers could demand additional evidence, including independent audits, sample tasks, error analysis and clear explanations of how scores were produced.

The most influential benchmark providers will not simply publish rankings. They will shape procurement decisions and, indirectly, the language companies use with investors and regulators. A favorable evaluation could support a model’s commercial launch. A poor result could delay deployment or expose weaknesses that a public leaderboard never revealed.

Vals’ funding round is therefore significant for more than the company itself. It signals that evaluation is becoming a strategic layer of the AI industry, positioned between model development and adoption. The central question is no longer only which system is smartest. It is which system can be trusted to perform useful work, under realistic conditions, without creating risks that outweigh its benefits.

If private, industry specific testing earns that trust, it could become the infrastructure through which AI systems enter the economy. If it does not, buyers may face a new problem: benchmarks that are more realistic than public tests, but too opaque to serve as a dependable standard.

#Vals#Rayan Krishnan#Andreessen Horowitz#8VC#Bloomberg Beta#Geneva Convention
Alex Carter is an AI and technology journalist focused on how artificial intelligence is reshaping business, software, and everyday decision-making. He covers emerging models, industry shifts, and real-world adoption with an emphasis on what matters beyond the announcement.

This article was written with the assistance of an AI system and published automatically.