Imagine hiring an AI agent to solve a problem, then discovering that its performance depends less on its underlying model than on the memory, tools, supervision and recovery systems surrounding it. That is the future suggested by Nvidia’s latest agent benchmark result, where Claude Opus 5 moved from a roughly 30% score to 100% when placed inside a custom software harness.
In a report on the result, TechCrunch described Nvidia’s harness research as evidence that the supporting system around an AI model may be becoming more important than the model itself. The finding challenges the way businesses often compare AI products today, by placing Claude, GPT and Gemini on a leaderboard and treating the highest score as the obvious choice.
The more consequential question may be different: What happens when each model is given a carefully designed working environment?
Nvidia’s result suggests that an AI model should not be viewed as a finished worker waiting for a prompt. It may be closer to an employee entering an office. The quality of that employee’s tools, filing system, manager, feedback process and ability to recover from mistakes can determine how much of their ability becomes useful in practice.
From 30% to a perfect score
Nvidia said in its announcement that its AVO agent system achieved a 100.00 RHAE score across all 25 public environments in the ARC-AGI-3 benchmark while using Claude Opus 5 as its underlying model.
The comparison is striking. The ARC Prize Foundation’s verified Claude Opus 5 results record the model at 30.16% on ARC-AGI-3 at high reasoning effort. Nvidia’s reported result therefore presents a model that moved from solving about three out of every ten benchmark environments to solving all of them when integrated into a larger agent architecture.
That does not mean Claude Opus 5 suddenly became a different model. The change was in the system surrounding it. Nvidia attributes the result to a combination of memory, tools, feedback, supervision and recovery mechanisms. Those components can alter how an agent approaches a task, how long it can work on it, how it notices errors and whether it can use a failed attempt to improve its next one.
The distinction matters because a model score is often presented as though it represents a stable property, similar to a processor’s speed or a camera’s resolution. Nvidia’s experiment suggests that agent performance may behave more like productivity in a workplace. The same person can produce very different results with poor information, limited tools and no useful feedback compared with an environment designed around their strengths.
What the harness actually does
Nvidia’s AVO system is not described as a simple prompt wrapper. In the researchers’ technical paper on AVO, they describe an agent loop built for autonomous evolutionary search and long-horizon work.
The system maintains persistent context, executes actions, receives feedback, critiques its own work, repairs failures and verifies results. These steps create a process around the model. Instead of asking Claude to solve a problem in one uninterrupted attempt, the harness lets the agent build a working record, test ideas, examine what happened and revise its approach.
That structure is particularly important for interactive tasks. An agent may need to inspect an environment, form a hypothesis, take an action, observe the outcome and then decide what to do next. If each step is isolated, an error can erase the value of previous work. If the system retains context and records useful discoveries, a mistake can become part of the solution.
A supervisory agent adds another layer. It can decide how to divide work, when to request another attempt and whether an answer appears sufficiently reliable. The result is less like a chatbot producing a response and more like a small team operating around a shared project.
This changes the source of capability. The model still supplies language, reasoning and pattern recognition. But the harness determines how those abilities are organized over time. It can transform a collection of promising individual answers into a repeatable problem-solving process.
Why ARC-AGI-3 exposes the difference
ARC-AGI-3 is designed around interaction rather than a single question followed by a single answer. The benchmark’s technical report explains that its evaluation involves interactive environments and that harness design, context management and external scaffolding can materially influence agent performance.
That makes the benchmark especially revealing for the current generation of AI systems. Many familiar tests ask a model to respond to a fixed prompt. Such evaluations are useful for comparing some aspects of knowledge and reasoning, but they provide only a narrow view of what happens when an agent must act repeatedly inside an unfamiliar environment.
In a long-horizon task, the agent may need to remember details from earlier attempts, manage limited context, select tools and determine whether progress is real. It must also recognize when it is stuck. Those are not merely model abilities. They are system design problems.
A benchmark that includes these conditions therefore measures more than the intelligence of a model in isolation. It measures the intelligence of the arrangement. A weaker model with strong memory and recovery procedures may outperform a stronger model that is repeatedly forced to start over. A supervisory layer may spot a bad strategy before the main agent spends too much time pursuing it.
The implication is not that the model no longer matters. A system still needs a capable engine. Rather, the result suggests that model quality is only one part of the performance curve, and perhaps not the part with the greatest room for improvement.
The end of the simple leaderboard
For enterprise buyers, this could make today’s model comparisons much less useful.
A procurement team might currently compare benchmark tables, pricing and context windows before choosing a model. But if the same underlying model can perform radically differently under different harnesses, a score attached only to the model may conceal the most important information.
The relevant comparison could become a full stack: model, memory architecture, tool permissions, orchestration logic, supervisor, recovery strategy, evaluation process and operating cost. Two vendors might offer access to the same base model while delivering very different results because one has built a better environment for using it.
This would resemble the difference between buying a powerful engine and buying a reliable vehicle. The engine matters, but so do the transmission, controls, navigation system, maintenance procedures and safety features. A benchmark that reports only engine power cannot tell a buyer which vehicle will complete a difficult journey.
The shift could also complicate claims made by AI vendors. A company may advertise that its model performs better on a task, while the real advantage comes from a proprietary harness, undisclosed prompting strategy or extensive external tooling. Conversely, a model may appear weaker because it was evaluated with a minimal setup while a competitor received a more sophisticated agent layer.
That does not automatically make either result misleading. It does mean the evaluation needs to describe the complete conditions under which the result was produced.
Harnesses as a new competitive moat
If the value moves from the model alone to the entire agent system, software harnesses could become a major source of competitive advantage.
Models are increasingly accessed through application programming interfaces, which means several companies can build products on top of similar underlying systems. The differentiator may be the invisible machinery that turns a general-purpose model into a dependable worker. That machinery can include specialized memory, task planning, tool selection, error detection and domain-specific verification.
For a company building an AI research assistant, the key advantage may not be a slightly stronger model. It may be a system that remembers which sources were checked, tracks unresolved questions, flags contradictions and knows when to ask a human for help. For an engineering agent, it might be the ability to run tests, inspect failures, patch code and verify that a fix did not create a new problem.
These capabilities are difficult to judge from a short demonstration. They emerge over dozens or hundreds of actions, especially when conditions are messy. That makes harness engineering less visible than a model launch, but potentially more important to the user’s daily experience.
The result could be a new layer of AI infrastructure companies competing to build the best operating environments for agents. Model providers may respond by offering their own orchestration tools, memory systems and supervisory agents. Application companies may keep the model interchangeable while protecting their harness as their core intellectual property.
What companies should ask instead
Nvidia’s result does not prove that every custom harness will deliver a dramatic improvement. It does show why organizations should ask more detailed questions before treating a benchmark score as a buying decision.
They should ask whether a result is model-only or system-level. They should examine the number of attempts permitted, the tools available, how context is stored and whether human intervention is allowed. They should also ask how failures are handled, whether the agent can retry with a different strategy and how the final result is verified.
Cost matters as well. A system that reaches a perfect benchmark score by running many agents, repeated searches and extensive verification may be more expensive and slower than a simpler system. In a business setting, a lower score could be preferable if it arrives quickly, uses fewer resources and is easier to audit.
Reliability should be measured across ordinary tasks, not only difficult benchmark environments. An agent that succeeds spectacularly in a controlled test but loses track of customer requirements in routine work may not be useful. Enterprises will need evaluations that combine capability with consistency, transparency, latency, security and the ability to hand work back to people.
A more honest picture of AI performance
The most important consequence of Nvidia’s finding may be cultural. It encourages the industry to stop talking about AI models as if they operate alone.
In the near future, people may rarely interact with a raw model. They will interact with an assembled system that chooses tools, stores memories, asks for confirmation, delegates subtasks and checks its own output. The user will experience the harness as the product, even if the model remains the most celebrated component.
That could produce better AI systems. Companies would compete not only to build larger models, but also to make them more dependable, understandable and useful in real workflows. Designers would have to consider how an agent communicates uncertainty, how it recovers from a wrong turn and how much control a person retains.
It could also produce more honest benchmarks. Instead of displaying one number beside a model name, evaluations may report the complete agent stack and the conditions behind the result. A score would then tell buyers not simply what a model knows, but what a working system can accomplish.
Nvidia’s AVO result is therefore more than an impressive benchmark jump. It is a preview of an AI market in which the decisive innovation may sit around the model rather than inside it. As agents move from answering questions to carrying out extended work, the quality of their memory, tools, supervision and feedback loops may determine whether they feel like clever demonstrations or dependable collaborators.
- Coolcaesar · CC BY-SA 4.0
This article was generated using AI and published automatically without human pre-publication review.
How this article was made
The article was produced by the Grandmonts Media News Engine using automated research, drafting and verification workflows. No human editor reviewed the article before publication. Grandmonts Media remains responsible for the published content. Errors can be reported at office@grandmonts.cz.