A new research agent from DeepMind alumni is challenging the assumption that scientific AI needs the largest possible model. Inherent says its Faraday system, built around a 27 billion parameter model, reproduced published research more successfully than much larger systems from OpenAI and Anthropic. The result is promising, but it also raises difficult questions about testing, human oversight and what it really means for an AI to do science.

For a scientist trying to reproduce someone else’s work, the hardest part is often not understanding the paper. It is discovering what the paper leaves unsaid.

A research article may describe an experiment in several pages, yet depend on dozens of small decisions that never make it into the final text. Which software version was used? How were failed trials handled? Which result looked promising enough to investigate further? What did the researchers try before arriving at the method that eventually appeared in print?

These details can determine whether another team succeeds or spends weeks chasing an apparent contradiction. They also represent a kind of judgment that is difficult to reduce to a checklist.

London startup Inherent is betting that artificial intelligence can learn some of that judgment. The company, founded by four former Google DeepMind employees, has emerged from stealth with an AI research agent called Faraday. Inherent says Faraday performed better than Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 on a task focused on independently reproducing results from published scientific papers.

The claim is notable for two reasons. First, replication is a practical test of whether an AI system can move beyond producing convincing text and actually complete a sequence of experiments. Second, Faraday is built around Qwen 3.6, a model with 27 billion parameters. That is far smaller than the frontier systems used in the comparison, according to Inherent.

The announcement therefore touches a growing fault line in the AI industry. One side continues to invest billions of dollars in larger models, bigger data centers and increasingly expensive training runs. The other is looking for ways to make smaller models more capable through reinforcement learning, careful tool use and systems that are designed around specific jobs.

If Inherent’s results hold up under independent examination, they could suggest that the next advances in scientific AI will not come only from scaling a general purpose model. They may come from teaching an agent how to work.

The science of choosing what to try

Inherent co-founder and chief scientist Edward Hughes describes the quality the company wants to develop as “research taste.” The phrase refers to the ability to select worthwhile experiments and create useful experimental designs, rather than simply follow instructions.

That distinction matters because scientific work is full of decisions that are not fully specified in advance. A researcher must decide which hypothesis is worth testing, which variables are most important, how to interpret an ambiguous result and whether a failed experiment reflects a bad idea or a flawed implementation.

A system that only follows a written plan may be useful for routine laboratory or coding tasks. A system that can recognize when a plan is unproductive, revise it and try a better approach would be closer to a research collaborator.

Inherent says its training approach relies heavily on reinforcement learning. Rather than attempting to encode every element of good scientific practice as a fixed set of rules, the company rewards the system for producing useful outcomes. In principle, that allows Faraday to learn which choices tend to lead to successful results, even when the precise path differs from one paper to another.

The concept resembles the way a junior researcher develops expertise. At first, the researcher may need explicit guidance for each stage of a project. Over time, repeated experiments and feedback build an intuition about where to spend effort. The researcher learns not just how to run an experiment, but which experiment is worth running next.

That intuition is difficult to measure, however. A system can appear to possess scientific judgment while simply exploiting weaknesses in a benchmark. It might recognize patterns in the evaluation set, rely on hidden assumptions about common research methods or benefit from generous scoring rules.

The quality of Inherent’s evidence will therefore depend less on the label “research taste” than on the details of how Faraday was tested.

Replication is not discovery

Reproducing an existing scientific result is a meaningful challenge, but it is not the same as making an original discovery.

Replication asks whether an agent can take a published claim, reconstruct the relevant method and obtain a result that is sufficiently similar. This requires reading comprehension, planning, coding, debugging and persistence. It can also reveal whether an AI system is able to cope with incomplete documentation and unexpected failures.

But the target is already known. The agent has a destination, even if the route is not provided.

Original research is more open ended. It requires deciding which questions deserve attention, identifying patterns that have not been reported, proposing explanations and determining whether a surprising observation is a genuine insight or an artifact. There is no guaranteed answer against which every step can be checked.

That difference does not make replication unimportant. Scientific progress depends on the ability of other researchers to verify and extend published work. In many fields, replication is slow and expensive because papers omit implementation details, research code is difficult to use or the original team made undocumented choices.

An agent that can reliably reproduce results could reduce some of that friction. It might help laboratories audit claims, identify fragile findings and test whether a proposed method works in a different environment. It could also allow researchers to examine a larger number of papers than a human team could handle alone.

Still, a strong replication benchmark should be described as a test of research automation, not proof that an AI scientist has arrived. The distinction matters because the language used to describe these systems can shape investment decisions, public expectations and the way researchers evaluate their own work.

The importance of the surrounding system

Faraday’s reported performance also complicates the usual argument about model size. The central language model may be relatively small, but the agent is not operating in isolation.

For coding tasks, Inherent had Faraday use OpenAI’s GPT-5.5 Codex. That means the system combines different components rather than asking one model to perform every task. Faraday may handle planning, experiment selection and broader research decisions, while a specialized coding system helps turn those decisions into executable programs.

This approach is closer to assembling a research team than choosing a single employee. One member may be strong at forming hypotheses, another at writing code and another at checking results. The effectiveness of the group depends on how responsibilities are divided and how information moves between them.

The distinction is important when comparing systems. If one model is evaluated as a standalone assistant while another can call additional tools, use external models or repeat experiments over a longer period, the result reflects more than the underlying model. It reflects the architecture, budget and operating rules of the entire system.

A smaller model can outperform a larger one in a narrow setting if it has been trained specifically for the task. A chef with a limited kitchen may produce a better meal than a more famous chef if the ingredients, recipe and time constraints favor the first. That does not establish that the first chef is better at every kind of cooking.

The same logic applies to AI agents. Faraday’s result, if confirmed, would show that a carefully designed research workflow can overcome some disadvantages in raw model scale. It would not necessarily show that a 27 billion parameter model is broadly more capable than the systems it beat on this test.

The benchmark questions that matter

Inherent’s comparison is self-reported, making the structure of the evaluation central to the credibility of the claim.

The first question is what papers were selected. Were they drawn from a broad collection of fields, or chosen because their methods were especially compatible with Faraday’s training? Were the papers published before the systems were developed? If the agent had encountered similar examples during training, its apparent independence could be overstated.

The second question concerns the information given to each system. Did Faraday and the competing models receive the same papers, instructions, computing resources and access to tools? Were they allowed to search the internet or consult additional documentation? Could they ask for human help when they became stuck?

The third question is how success was measured. A replication can be considered successful under different standards. The agent might need to match a numerical result, recreate a trend, reproduce a statistical conclusion or simply implement a method that runs without errors. Each standard measures something different.

Runtime and cost are also essential. An agent that succeeds after hundreds of experiments and many hours of computation may be less practical than one that reaches a slightly weaker result quickly. Conversely, an expensive workflow could still be valuable if it replaces months of work by a research team.

Human intervention must be reported with equal care. There is a major difference between a system that works independently from the paper and a system that receives occasional hints from an engineer. Both can be useful, but they represent different levels of automation.

Finally, other researchers need to be able to reproduce the comparison. That requires access to the evaluation papers, prompts, model settings, tool permissions, scoring rules and relevant logs. Without those materials, outsiders may be able to admire the result but not meaningfully test it.

A new competitive landscape for AI models

The launch arrives as the AI industry searches for ways to make increasingly capable systems economically sustainable. Frontier models require extraordinary investments in chips, data centers and energy. Their generality is valuable, but the cost of training and operating them can be difficult to justify for specialized work.

Smaller models offer a different path. They may be cheaper to run, easier to deploy and more suitable for organizations that cannot afford constant access to the largest commercial systems. Open weight models can also give companies and researchers more control over privacy, customization and infrastructure.

The risk is that smaller systems can look stronger than they are when judged only on carefully selected tasks. A model may perform impressively in a constrained research workflow while struggling with unfamiliar domains, weakly documented papers or experiments that require physical laboratory equipment.

The more durable opportunity may lie in hybrid systems. Instead of asking one general model to manage every part of a project, companies can combine models with different strengths, software tools, retrieval systems, simulators and human reviewers. The result could be less like an autonomous genius and more like an always available research operations team.

That would still represent a significant change in how scientific work is organized. Researchers could delegate literature review, replication attempts and routine computational experiments to agents, then spend more time interpreting results and choosing broader directions.

But delegation creates responsibility. If an agent produces a misleading replication, who is accountable for the resulting paper or decision? If an AI system prioritizes experiments according to a reward function, what important questions might it ignore because they are difficult to score? And if research organizations begin to rely on automated systems, will scientists lose opportunities to develop the judgment that the systems are meant to imitate?

What Faraday may prove

The most useful interpretation of Inherent’s announcement is neither that model scale no longer matters nor that AI has already become an independent scientist.

Faraday may demonstrate that performance in scientific work depends on a broader recipe than parameter count. Training methods, experiment selection, persistence, software access and evaluation design can materially change what a model is able to accomplish. That is an important lesson for companies that have treated the largest model as the inevitable foundation for every application.

The result may also encourage more serious evaluation of AI research agents. Instead of judging them by polished demonstrations or generated summaries, developers can ask whether they complete verifiable tasks, recover from failure and produce results that other people can inspect.

Replication is a particularly useful place to begin because it creates a bridge between language and evidence. The agent must not only say what a paper means. It must act on that understanding and face the consequences when its interpretation is wrong.

The harder test will come later. Can Faraday reproduce work from unfamiliar fields? Can it distinguish a promising anomaly from a coding error? Can it design an experiment whose value was not already implied by the paper it was given? Can independent researchers obtain the same results without relying on Inherent’s own infrastructure?

Those questions will determine whether the system is a narrow but effective automation tool or an early example of a more general scientific collaborator.

For now, Inherent’s announcement points to a shift in the center of gravity of AI competition. The decisive advantage may not always belong to the company with the biggest model. It may belong to the team that understands how to turn a model into a disciplined partner, gives it the right tools and measures its work against reality.

That is a more modest claim than the arrival of an artificial scientist. It may also be a more consequential one. Science advances through thousands of choices about what to test next. If AI can help make those choices more efficiently, even within the limited world of replication, it could begin changing research long before it is capable of making discoveries on its own.

#Inherent#Faraday#Edward Hughes#Qwen 3.6#GPT-5.5#Claude Opus 4.8#Google DeepMind
About Daniel Reyes
Daniel Reyes writes spAIsee's technical explainers: how a model is built, trained, evaluated and served, and where the published claims stop matching the measured behaviour. He covers architecture, inference economics, evaluation methodology and agent tooling, and reads the paper before the press release.