DTWdailytechwire
Tech Intelligence, Wired Daily
AI

A 27-Billion-Parameter Model Just Beat Frontier Labs at Scientific Paper Replication

London's Inherent built an AI agent that outperforms far larger systems by embedding research taste through reinforcement learning, not scale

DR
Daniel R. Whitfield
Markets & Venture Reporter · Hong Kong
Aug 23, 2026
5 min read
A 27-Billion-Parameter Model Just Beat Frontier Labs at Scientific Paper Replication
A 27-Billion-Parameter Model Just Beat Frontier Labs at Scientific Paper ReplicationCredit: Anna Gordon

The Benchmark That Matters

A London AI lab has demonstrated that teaching machines to think like scientists may matter more than feeding them endless parameters. Inherent, a startup led by former researchers from a major UK AI institution, released results showing its agent Faraday successfully replicated findings from published scientific papers using a model containing just 27 billion parameters. The same task, when attempted by systems built on architectures ten times larger, produced weaker results.

The test itself is straightforward: given a scientific paper, can an AI independently reproduce its experimental conclusions without being handed the answer upfront? For human researchers, particularly doctoral candidates, this exercise serves as a rite of passage. It builds intuition about which variables matter, how to structure experiments, and when to pivot. Inherent's thesis is that the same training ground applies to machines.

What separates this result from typical benchmark races is not the margin of victory but the method. Faraday runs on Qwen 3.6, a comparatively compact foundation model. Against it stood systems with parameter counts in the hundreds of billions, trained on budgets that would fund dozens of startups. The smaller model won not by brute-force memorization but by learning what the team calls research taste.

Reinforcement Learning as a Substitute for Scale

Research taste is an elusive quality. It encompasses knowing which experiments are worth running, recognizing when a hypothesis has been tested sufficiently, and understanding how to allocate limited time across competing priorities. These instincts are difficult to encode in supervised learning, where models learn by imitating labeled examples.

Inherent's approach leans heavily on reinforcement learning, a training paradigm that rewards agents for achieving desired outcomes rather than following prescribed steps. According to Edward Hughes, the company's chief scientist, this choice reflects a long-term bet. The startup aims to build AI capable of contributing original scientific insights across disciplines, not just validating existing work. Training agents to mimic the structure of published research would teach them the form of science, not its substance.

By structuring Faraday's training around reward signals tied to experimental success, Inherent hopes the agent will generalize beyond paper replication to hypothesis generation and experimental design in unfamiliar domains. The results so far suggest the strategy has merit. Faraday not only matched the accuracy of larger systems but also demonstrated judgment about which experiments to prioritize, a quality the team measures separately from raw correctness.

The computational efficiency is striking. Parameter count correlates loosely with training cost, inference latency, and hardware requirements. A 27-billion-parameter model can run on infrastructure orders of magnitude cheaper than frontier systems, making it feasible to deploy such agents in academic labs or small research teams that lack access to hyperscale compute clusters.

What Inherent Chose Not to Build

Equally revealing is what Inherent decided to outsource. Rather than developing a proprietary code-generation tool, the team configured Faraday to call existing coding models when it needs to write or debug software. The rationale mirrors how human scientists work: they reach for established tools rather than rebuilding the stack from scratch for every experiment.

This pragmatism extends to the company's broader philosophy. Hughes described the ideal AI teammate as one that returns from a task with unexpected findings, not just confirmation of what the user already believed. The goal is an agent that exhibits curiosity and initiative, qualities that require moving beyond instruction-following architectures.

Inherent's entire staff of twelve works in person from an office in King's Cross, a district in north London that has become a focal point for AI research talent. Hughes emphasized the value of geographic concentration, particularly in a city where competition for experienced researchers is intense. The company plans to expand to between twenty and twenty-five employees by year-end, a measured pace compared to the hiring surges seen at better-funded competitors.

The Garden Leave Problem

Hughes has publicly criticized a UK employment practice that complicates startup formation: garden leave clauses that prevent departing employees from joining or founding competing ventures for months after resignation. The restriction is standard in British contracts but rare in the United States, where researchers can move between organizations with minimal delay.

The asymmetry creates a structural disadvantage for UK startups competing for talent against American counterparts. A researcher who leaves a major lab in California can join a new venture almost immediately; a peer in London may wait half a year or longer. Hughes noted that he personally navigated this constraint before launching Inherent, though he characterized his criticism as a personal view rather than official company policy.

The issue has broader implications as London positions itself as a counterweight to Silicon Valley in AI development. The city benefits from a concentration of academic institutions and corporate research labs, but regulatory and contractual friction can slow the flow of talent into startups. If Inherent's hiring ambitions are any indication, the company sees an opportunity to attract researchers weighing their next move, particularly those unsettled by recent leadership changes at larger organizations.

Beyond Replication

Paper replication is a means, not an end. Inherent's north star remains building an AI scientist capable of generating novel hypotheses and conducting original research. The Faraday release serves as proof of concept for the underlying training methodology, but the team has not disclosed timelines for when it expects to demonstrate true scientific discovery.

The gap between replication and innovation is substantial. Reproducing published results requires following a known path; discovering new knowledge demands recognizing which paths have not yet been explored. The latter involves higher-order reasoning about what questions are worth asking, a challenge that has so far eluded even the largest language models.

Inherent's bet is that reinforcement learning, combined with smaller, more efficient architectures, will scale to this level of reasoning faster than approaches that rely on ever-expanding parameter counts. The company emerged from stealth weeks ago with fifty million dollars in seed funding, a signal that investors see merit in the thesis. Whether that capital proves sufficient to reach the goal of autonomous scientific discovery remains an open question.

At DailyTechWire, we have tracked the proliferation of AI labs promising breakthroughs in scientific research. Most have delivered little beyond vague roadmaps and models that excel at narrow benchmarks. Inherent's willingness to release a working agent and benchmark it against named competitors is a departure from that pattern. The results are preliminary, but they suggest a path to useful scientific AI that does not require the compute budgets of frontier labs.

The real test will come when Faraday or its successors move beyond replication to contribute findings that human researchers consider worth pursuing. Until then, the London team has at least demonstrated that smaller models, trained with care, can compete with giants.

Read next
AI

Nvidia Server Prices Jump as Memory Costs Squeeze AI Infrastructure Buyers

Arjun S. Mehta · 4 min
AI

Claude's Older Models Still Bypass Safety Rules on Sexual Content

Priya Nair · 6 min
AI

The Software Wrapper Around AI Models Now Outweighs the Model Itself

Arjun S. Mehta · 6 min
Spot something wrong? Email corrections@dailytechwire.com. We log every correction publicly.