Benchmarks

A new benchmark says frontier AI recovers a paper's idea 3% of the time. Read it before you decide what that proves.

'Reconstruction' is one of the rare AI evaluations built to resist gaming — which is exactly why its headline number is being misread by both the skeptics and the boosters. The story is in the error bars neither side will screenshot.

Old book spines — the accumulated research literature a new benchmark asks AI models to reason from

Image: Tom Murphy VII / Wikimedia Commons (CC BY-SA 3.0)

The number that will get screenshotted is 3 percent. A benchmark posted to arXiv on August 17 handed seven frontier models a research paper's reference list — and nothing else — and asked each to propose the idea the paper was actually about. The match rate came in between roughly 3 and 15 percent. Cue the takes already writing themselves: 'AI can't do science.' Before you repost that, three questions decide what the result is worth. Matched against what? Judged by whom? And 3 percent of exactly what? The answers make this finding more interesting than the headline — and less useful to both cheering squads than either would like.

The benchmark is called Reconstruction, and the design is elegant and a little cruel. Take a published paper, strip it down to its pre-publication bibliography — only the works it cited, with nothing from the paper itself and nothing published alongside or after it — and ask a model to reconstruct the central hypothesis the authors went on to make. An independent language-model judge then decides whether the model's proposal matches the real idea, which has been held out the entire time. The authors ran this across 643 papers in six scientific domains. It is, on its face, a clean test of whether a model can look at what a field had just read and infer what someone was about to figure out.

Why this test is harder to game than most

Start with the part I want to praise, because I spend most of my time doing the opposite. Most benchmarks rot, and they rot for a predictable reason: once a test is public and everyone optimises for it, it stops measuring capability and starts measuring exposure to the test. The scores climb; the meaning drains. Reconstruction was clearly built by people who know this in their bones. It ships with three anti-leakage measures that do the boring, unglamorous work of keeping the test honest — and that is precisely the work that usually gets skipped.

  • A temporal citation cutoff: the model only sees references that predate the paper, so it cannot 'cite forward' to work that would give the answer away.
  • Anonymous reference IDs: the bibliography is stripped of the identifying details that would let a model simply recognise which paper this is and recall its abstract from training.
  • Frozen per-paper bibliographies: each test item is fixed, so the reference set can't drift or be quietly enriched into a giveaway.

That is the difference between a measurement and a magic trick. A test that leaks the answer through the citations would produce a flattering number that means nothing; this one goes out of its way to shut those doors. When a benchmark is this careful about contamination, it earns the right to be taken seriously — which, in this field, is rarer than it should be. So credit where it is due, by name.

Contamination isn't fraud, mostly. It's gravity. Reconstruction is one of the few benchmarks built to fight the pull. — Phuong Nguyen

Compared to what?

Now the seam. The task is not 'generate a good scientific hypothesis.' It is 'reproduce the specific hypothesis these particular authors landed on, using only their citations.' Those are not the same test, and the gap between them is where the headline overreaches. Science rarely has a single correct next idea. A strong bibliography can point toward several defensible hypotheses, and Reconstruction scores a model against exactly one of them — the one that happened to get written up and published. A model could propose something genuinely sensible, even something arguably better, and still score a clean zero because it was not what the authors did. Read that way, 3 percent is not obviously a measure of how good the models are at ideas. It is at least partly a measure of how convergent the single target is.

Which is why the number I most want is the one I do not have: the human baseline. Hand a domain expert the same anonymised reference list and ask them to reconstruct the paper's idea. What do they score? If a specialist hits 80 percent, the models are genuinely, embarrassingly far behind. If a specialist scrapes 20 percent, then the task itself is close to impossible and a machine's 3 to 15 percent is far less damning than it reads. 'Compared to what' is not a rhetorical tic here. It is the entire difference between 'AI is bad at science' and 'this is a brutally hard reconstruction problem and everyone struggles at it.' Without the human line drawn on the same chart, a bare 3 percent is a number without a ruler.

Then there is the judge. Whether a proposed hypothesis 'matches' the held-out idea is a soft, semantic call — and here that call is made by another language model. As a way to score thousands of open-ended answers, that is defensible and scalable. It also carries its own error bars, and they belong in the report. How often does the judge wave through a plausible near-miss? How often does it fail a correct answer worded differently from the original? A benchmark this fastidious about leakage should be held to the same standard on its scoring, and the honest way to present a model-graded result is with the grader's own reliability stapled to it.

The 42 percent that isn't the story either

There is a second number in this paper, and it will be cherry-picked in the opposite direction. When the authors stopped asking one model and instead built a pipeline — models reviewing each other's guesses, with the competing hypotheses run through a bracket-style Swiss tournament — the match rate climbed to between 23 and 42 percent across the six domains, roughly a 2.4-times lift over the best single model. Somewhere a post is already being drafted that reads '42 percent' as proof the AI co-scientist has arrived. Read it honestly and the tournament result says two things at once, and you have to hold both.

The first: aggregating many model guesses and adversarially filtering them helps a great deal. That is a real and genuinely useful finding about how to deploy these systems — not as a single oracle, but as a population you make compete. The second: even after that 2.4-times lift, the pipeline still misses the majority of the ideas. Both halves are equally true, and only one of them fits in a triumphant repost. The correct summary is the deflating, accurate one: a tournament of models did meaningfully better than any single model, and still could not recover most of the answers.

A tournament of models did 2.4 times better than one — and still missed most of the answers. Screenshot both numbers, or neither.

What the number actually supports

So strip it back to what the evidence will carry, which is the only thing a data editor is ever really for. Given only a paper's citations, and rigorously denied any glimpse of the answer, current frontier models rarely reconstruct the one specific hypothesis that paper made: low single digits to the mid-teens alone, up into the low forties as a coordinated ensemble. That is a well-built, blind, hard-to-contaminate result about a genuinely difficult task, and it deserves to be read as exactly that — no more, no less.

What it is not, on its own, is a verdict that machines cannot originate scientific ideas. 'Reproduce this exact paper's hypothesis from its bibliography' is a proxy for ideation, and a strict, single-target one, reported without a human baseline and leaning on a model-graded notion of a match. It is strong evidence about retrieval, about how convergent scientific reasoning is, and about the surprising payoff of making models argue with each other. It is weak evidence about the ceiling on machine creativity, because it was never built to measure that ceiling. Keep those two claims apart and the paper becomes genuinely useful. Blur them and you get another viral number that means less than it says.

The reason to like this study is not that it embarrasses the labs. It is that it is the kind of measurement the field keeps failing to make: blind, hard to contaminate, honest about its own difficulty, and published so anyone can check the working. The reason to handle it with care is that both camps arrive with a number pre-loaded — 3 percent for the skeptics, 42 for the boosters — and the truth is sitting in the error bars that neither will screenshot. My standing rule is that when a chart only goes up, you look harder. This one mostly goes sideways, which is rarer, and more honest, than either headline is going to let it be.

References

  1. arXiv — Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies (2608.16645)
  2. arXiv (full text) — Reconstruction: A Blind Benchmark for Recovering Research Ideas
  3. TechTimes — Blind Benchmark Catches Frontier AI at Just Three Percent on Research Idea Recovery
  4. arXiv — AI Idea Bench 2025: AI Research Idea Generation Benchmark (2504.14191)
  5. Stanford HAI — 2026 AI Index Report: Technical Performance
The Friday Brief

One email. Every Friday.

The week's machines, money, and people — in under five minutes.