Salesforce researchers say DarwinX lifted a browser agent’s WebArena-Infinity score from 43.5% to 93% without changing the underlying model, pointing to a new competitive battleground for enterprise AI: the software layer that controls how agents work.

The model is only part of the product

The result, reported by VentureBeat, challenges the assumption that meaningful agent progress depends primarily on the next frontier model. Salesforce AI Research and Salesforce Agentforce developed DarwinX to improve an agent’s “harness,” the surrounding system of prompts, tools, skills, permissions, control flow and workflow logic.

DarwinX’s reported benchmark gains, includingWebArena-Infinity rising from 43.5% to 93%.percentage point02040SWE-bench Verified3.4WebArena-Infinity49.5
DarwinX’s reported benchmark gains, including WebArena-Infinity rising from 43.5% to 93%.

That distinction matters commercially. Most companies cannot retrain the leading language models, and many cannot inspect or modify their weights. They can, however, decide which tools an agent may use, how it handles errors, when it asks for approval and how it breaks a complex task into steps. If those controls become the main source of performance gains, the companies that build better harnesses could create durable advantages even when they rely on the same underlying model as their rivals.

Salesforce says DarwinX improved performance across four benchmarks. The largest reported gain came on WebArena-Infinity, where the score rose from 43.5% to 93% with GPT-5.5 held fixed. The reported gains ranged from 3.4 percentage points on SWE-bench Verified to 49.5 points on WebArena-Infinity.

Salesforce Tower from Salesforce Park
Salesforce Tower from Salesforce Park · Dead.rabbit · via wikipedia · CC BY-SA 4.0

The researchers also said they conducted an audit to screen out invalid or exploit-style behavior. That qualification is important. An agent that achieves a benchmark objective by exploiting a test weakness may look more capable without becoming more useful or trustworthy in production.

An alternative to endless self-editing

DarwinX is designed around a problem Salesforce says affects common self-improvement systems: path dependence. When a system continually edits a single incumbent configuration, an early change can determine the direction of all later changes. A modification that solves one failure may push the agent toward a local optimum or weaken its performance on another task.

Cross-task interference creates a similar problem. Improving browser navigation, for example, could damage coding or planning behavior if the same instructions and workflow logic are repeatedly adjusted without sufficient safeguards.

Rather than maintaining one evolving version, DarwinX generates multiple harness variants, preserves an archive of promising alternatives and advances changes only when they improve results without unacceptable regressions. The approach resembles evolutionary search, but its business value lies in controlled iteration. It could allow enterprises to improve agents without repeatedly risking the reliability of their existing workflows.

Salesforce also reported some transfer between tasks. A harness evolved on Terminal-Bench improved performance when applied unchanged to SWE-bench Verified. The evidence remains limited because the transfer experiment was tested in only one direction, but it suggests that some harness improvements may capture general operating principles rather than benchmark-specific tricks.

Evaluation becomes the strategic bottleneck

The commercial opportunity is substantial. Better harnesses could reduce the cost of deploying agents, extend the useful life of existing models and give Salesforce a stronger position against competitors such as Microsoft, Google and independent agent platforms. The moat would not be the model alone. It would be the accumulated system of evaluations, workflow patterns and operational safeguards surrounding that model.

The constraint is equally significant. A company cannot safely automate a system that changes its own behavior unless it can detect subtle regressions. Enterprise workflows involve ambiguous requests, incomplete data, shifting permissions and consequences that benchmarks rarely capture.

DarwinX therefore points to a less glamorous but more important investment priority: evaluation infrastructure. The winners in enterprise agents may be the companies that can measure reliability across real business scenarios, not simply those that produce the highest benchmark score. Harness engineering may become the competitive layer, but trustworthy testing will determine whether that layer can support production automation.

#Salesforce#DarwinX#Agentforce#GPT-5.5#WebArena-Infinity#SWE-bench Verified#Terminal-Bench
Image credits
Rebeca Smith is an AI and technology journalist specializing in the business of artificial intelligence. Her reporting focuses on the companies, investments, and competitive strategies driving the industry's rapid evolution. She closely follows Big Tech, AI startups, venture capital, semiconductor manufacturers, and enterprise software, explaining how commercial decisions shape the future of AI adoption. Rebeca's work combines financial insight with technological understanding, helping readers see beyond product launches to the economic forces transforming the industry.

This article was written with the assistance of an AI system and published automatically.