A Small Agent Against Two Frontier Giants
Inherent, a London AI lab founded by former DeepMind researchers, said its agent Faraday outperformed both Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 at independently reproducing the findings of published scientific papers.
Faraday is built on Qwen 3.6, an open-weight base model with 27 billion parameters, far smaller than the frontier-scale systems it was tested against. The task was not simple question answering: researchers gave Faraday a published paper's setup and asked it to reproduce the result without being told the answer in advance, then compared its output against the two much larger models on the same papers. Inherent published the methodology and results on its own research page, and the comparison was independently reported by TechCrunch.
The result lands just weeks after Inherent emerged from stealth with a 50 million dollar seed round, an unusually large seed for a European AI lab and a bet that a small, well-trained agent can compete with systems built by companies spending far more on compute.
Why It Won: Taste, Not Just Scale
The mechanism: Inherent says Faraday was trained to develop "research taste," an instinct for which experiments are worth running and how to design them well, using reinforcement learning that rewards good outcomes rather than a rulebook of correct steps.
That approach targets a narrower skill than general-purpose reasoning: knowing how to replicate one specific paper's finding is a bounded task, and a smaller model trained specifically for it can plausibly out-perform a general frontier model that was never optimized for that exact job. This lands against a backdrop of rising frontier-compute costs: Nvidia has told major cloud customers that AI server prices for 2027 delivery are rising more than 15 percent, driven by memory costs, which raises the price of defaulting to the largest available model for every task.
What It Means for AI Buyers
The bottom line: a specialist agent a fraction of the size of a frontier model beating that frontier model at a defined task is evidence that model size is not the only lever worth pulling when European and UK businesses budget for AI capability.
Procurement and technical teams evaluating AI vendors for a narrow, well-defined task, from document review to code migration to research support, now have a concrete reason to benchmark smaller, purpose-trained models against frontier general-purpose ones before committing to the larger and more expensive option. The frontier labs remain ahead on broad, general-purpose reasoning; the claim here is narrower and task-specific, not a claim that Faraday replaces Claude or GPT-5.5 across the board.
The Takeaway
A UK lab beating two well-funded US frontier labs at a defined scientific task, six months after founding, is a genuine data point for European AI competitiveness, and a reminder that the compute arms race is not the only race that matters.
Read next: Qwen's 3 Billion Downloads Hide a Compliance Gap | Alibaba Withheld the Number That Sets Your Cost



