Text version of this article
Also available as Markdown.
Science in Broad Daylight
AI agents are the first researchers to show all their work, and we still grade them on the answer.
A field note · October 2026
Chengyang Shi · Xianglin Ji · Jintao Huang · Jicheng Wang · Yifeng He · Jiachen Liu · October 2026
[Figure: Ink and watercolor: a little round-headed robot at a school desk holds up its exam sheet. The tiny answer box sits at the bottom, but the sheet unfolds into a long strip of handwritten working that spills off the desk onto the floor, and a human hand reaches in with a magnifying glass.]
In 1974 a project at Cambridge set out to publish every surviving letter Charles Darwin wrote or received. It found more than 15,000. Over the next 49 years more than ninety people worked on it, and the thirtieth and last volume came out in 2023. Those letters are the closest we can get to watching Darwin work. In them he tries out half-formed ideas on friends and reports on experiments still in progress.
Darwin is the rare exception. For most research the paper is all that survives, and the paper shows the one path that worked. A lab notebook, however carefully kept, misses most of what the researcher was thinking. Reviewers end up checking little more than the answer.
AI agents are the first researchers to show all their work. Every message they write, every command they run and every result they read back stays in the log. Nothing like this has existed before in the history of science. Every run comes with a complete account of how its research was done, open for anyone to check. Yet we still judge these agents by their final score, which amounts to grading the answer.
Science has always depended on witnesses. When the Royal Society met in the 1660s, experiments were performed in front of its members, and the reports named who had been in the room. Peer review grew out of the same idea. All of it was built for human researchers, who can only hand in answers, so all of it checks answers. Now there are researchers whose every step can be checked, and they produce research far faster than peer review can absorb. Some of the witnessing will have to be done by machines.
Open-Endedness Bench is our attempt at that witness. It is a benchmark that reads an agent’s full log, checks whether each conclusion rests on a result that actually ran, and scores the research process from what it finds. This post explains how it works and what it found in more than a hundred AI research runs.
The score
§1 What the Final Score Misses
Many of the research tasks we now give agents are open-ended. Nobody knows the best possible answer, so there is no answer key to grade against, and the only thing left to judge is the number the agent ends with.
That number hides a lot. Ten hours without progress might mean an agent is carefully ruling out dead ends, or it might mean the agent is flailing, and the score looks the same either way. A good score can hide something worse. In RE-Bench, one agent handed in a prediction of how performance grows with compute without ever training a model at a second compute budget.
Sometimes the score even rewards a mistake. Our favorite example comes from a race to train GPT-2 faster than the current record. Two hours and 48 minutes in, Claude Opus 5 changed one optimizer setting, Adam’s β₂, from 0.95 to 0.99, and the loss dropped from 3.30855 to 3.30777. It was pleased with itself:
First improvement: β2=0.99 → 3.30777 (−0.0008). Let me push further.Claude Opus 5, nanoGPT speedrun, after experiment 17
β2=0.99 is the optimum (0.998 is worse). Adopting it as the new base (3.30777).after experiment 18
It kept the new setting in the next 27 experiments it launched. A little later it designed a new optimizer, which it described as “the matrix analogue of the β₂ win.” With that optimizer switched on, the loss went up to 3.347.
Later in the same run, Opus 5 ran one configuration eight times without changing anything. We lined the eight results up the way the agent reads results, each against the best one before it.
Nothing changed, and it still “improved” twice
Same code, run 8 times. Nothing was changed. Each result is compared with the best one before it.
lossbest beforehow much better
run 13.27696
run 23.277703.27696
run 33.277773.27696
run 43.276313.276960.00065new best nothing changed
run 53.276433.27631
run 63.277203.27631
run 73.277043.27631
run 83.276223.276310.00009new best nothing changed
The change Opus 5 actually made, and what it wrote
β₂ 0.993.307773.308550.00078new best it wrote “First improvement!”
Losses from Opus 5’s speedrun, lower is better. The eight reruns are listed in the order its log reports them and trained for 3,150 steps; the β₂ change was measured at 2,700 steps against the best result so far at that length.
Identical code set a “new best” twice. The larger of those two gaps, 0.00065, is about the size of the β₂ gain the agent went on to build on, so there is no way to tell from this run whether β₂ = 0.99 helped at all. Opus 5 couldn’t tell either. (It set a new record eight hours later anyway, and the record is all the leaderboard keeps.)
Careful human scientists have made the same mistake. The best-known case is N-rays. In 1903 the French physicist René Blondlot announced a new kind of radiation, and within two years hundreds of papers had been written about it. In 1904 the American physicist Robert Wood went to Nancy to watch a demonstration. The room was kept dark, and while Blondlot read out lines of a spectrum, Wood reached into the apparatus and took out the prism.^1 Blondlot went on reading out the same lines. As far as anyone can tell, he believed every number. Opus 5’s β₂ win is probably the same kind of thing. The score it reported can’t tell you whether the change did anything. You find out only by going back to the trace and looking at what the agent ran and what came back, which is what the next section is about.
The method
§2 How We Read a Trace
Getting a machine to read traces fairly turned out to be harder than we expected. The first problem is that there is no answer key. The benchmark has to judge whether a claim follows from the evidence without knowing what the right answer is. The agent’s own account is no help here, since an agent can write “I ran the ablation” when no ablation ever ran. We also wanted one method for every task, because writing a new rulebook for each task would never end. And the reader makes mistakes too. We use a language model to read the traces, and our early versions went wrong in ways we hadn’t seen coming.
We got out of this by holding to one rule at every stage of the design.
Words can propose a claim. Only evidence returned by an action that actually ran can support it or refute it.
Wood did much the same in Nancy, when he ignored the readings and went to check the apparatus. Open-Endedness Bench does this in three steps. The first step is to convert each benchmark’s logs into one shared record, in which every step is marked as something the agent said, ran or saw. This conversion is the only code we write separately for each benchmark. Next, a language model reads the record a few steps at a time and writes cards, one kind for what the agent claimed and one for what it did. Every card has to carry an exact quote, and code throws away any card whose quote doesn’t appear word for word in the record. In the run below, seven cards went out this way.
The last step is done by ordinary code. It links the cards, noting for instance that this experiment tests that prediction and this result contradicts it, and every score is a count over the graph that comes out. We kept the language model’s part small on purpose. It copies quotes and answers yes-or-no questions, three times each with the majority winning, and ordinary code does everything else, including the scoring.^5
Take one of the runs we scored. Claude Opus 4.8 was given one GPU and ten hours to train a small language model to solve competition math problems. After its first round of training, the small model got 0 of 30 problems right.
Opus 4.8 found that the small model kept repeating itself. Its guess was that the trouble lay in how the small model picks each next word. By default it always takes the word it rates most likely, and that habit can trap a model in a loop. Opus 4.8 predicted that letting it pick with a little randomness would break the loops, so it made that change and had the small model take the same 30 problems again.
One answer came back right. Twenty-eight of the others still ran on until they were cut off at the maximum length. The change had barely helped, and the result went against the prediction. Opus 4.8 moved on to a new explanation, that the small model was “badly undertrained”, meaning it hadn’t been trained nearly enough, and it launched a much larger training run.
Here is that stretch of the log in the agent’s own words. The number on the left is the step, and the label next to it says whether the agent said something, ran a command, or saw what the command returned. The highlighted spans are the quotes that went onto cards.
199 saidDiagnosis: greedy decoding causes repetition loops. [...] Let me switch sft1_causal to temperature 0.6 / top_p 0.95 and re-eval. 203 ranpython evaluate.py --model-path out/sft1_causal --json-output-file results/sft1c_t06.json 212 saw"accuracy": 0.0333 · correct: 1/30 214 saidTemp 0.6 helped marginally (3.3%, 1/30) but 28/30 still ramble to 16k tokens [...] the model is badly undertrained and rarely concludes. 220 ranpython train_sft.py --out out/sft2 [...] --max-examples 12000 [...]
Steps 199 to 220 of the Opus 4.8 run, cut with [...] where the text is long.
When Open-Endedness Bench reads this stretch, it links the 1-in-30 result to the prediction it went against, then looks at how Opus 4.8 responded. The scoring code knows nothing about math problems or language models. It only knows what was said, what ran and what came back, which is why the same code can read any benchmark once the logs are converted. The first half of our demo follows this run from its raw log to the linked cards.
One real run, from its raw log to a shared record to cards and links. Thirty seconds.
The measures
§3 What We Count
There is more than one good way to do research. In the 1980s the psychologists David Klahr and Kevin Dunbar watched people work out how an unfamiliar device behaved. Some started with a theory and designed experiments to test it. Others ran experiments first and let a theory grow out of what they saw. Both groups got there. We took that seriously, so we describe an agent on two separate layers.
The first is competence, how carefully the agent handles evidence, and here higher is better. We measure it four ways. The evidence score checks whether the agent’s beliefs rest on results it has actually seen. The experiment score goes through every experiment the agent ran and asks whether it read the result and said what the result shows. Revision covers the moments when a result contradicts one of the agent’s ideas; the agent gets credit if it repairs the idea or explains the result. The last one, no reward hacking, checks that it leaves the grading alone.
The second is persona, the agent’s research habits: how long it thinks before acting, say, or whether it keeps polishing one idea or keeps reaching for new ones. Each habit is a scale between two ways of working, and good scientists use both. Neither end scores higher.
Most competence scores are built the same way. We count the chances the agent had to do the right thing, and how many of them it took. Every experiment, for instance, is a chance to read the result and say what it shows. The second half of the demo counts all 62 experiments in the Opus 4.8 run this way, then shows every score the run received. It ends with six models side by side, each matched with the scientist whose research habits it shares.
All 62 experiments in the Opus 4.8 run, sorted by what the agent did with each result, then every score the run received, then six models and the scientist each one works like, read from their post-training runs. Thirty-eight seconds. Portraits: Wikimedia Commons, public domain.
The runs
§4 What We Found in 119 Runs
We scored 119 runs that other people had already recorded, from three benchmarks. We didn’t run any new agents.^2 Four things stood out to us, and a final score shows none of them.
Most experiments never turn into a conclusion
The Opus 4.8 run above kept a conclusion from 30 of its 62 experiments, which turns out to be about average. Across all the post-training runs, 37 percent of experiments ended in a conclusion the agent kept. On chip design it was 24 percent. The rest were never read, read without comment, or taken back later. The speedrun was the odd one out at 85 percent, and we don’t have a good explanation for that yet. An agent that keeps a conclusion from only one experiment in three spends much of its budget on experiments it never learns from, and its final score gives no hint of that.
Most claimed wins are noise
The β₂ story turned out to be typical. On the speedrun and on chip design, the benchmark logged the true result of every experiment, so we could check each thing an agent said about its own results.^3
| When the agent said the result was… | Speed race | Chip design |
| an improvement | 16% | 29% |
| no improvement | 1% | 20% |
| (it said nothing) | 3% | 31% |
Share of experiments that truly beat the run’s best result so far, by what the agent said about them.
On the speedrun, the agents’ verdicts do mean something: a result they called an improvement was real more often than one they dismissed. They just called improvements far too often, 64 times for 10 real ones, because rerunning the same code moves the loss more than most changes do. On chip design we couldn’t find any signal. A result the agent called an improvement was real 29 percent of the time, and one it said nothing about, 31 percent.
This matters because agents build on their own verdicts. On the speedrun, 91 percent of the experiments an agent called improvements became the starting point for later experiments, and 60 percent of all follow-up experiments were built on a win that never happened.
The best runs keep exploring, and stop later
For each task we compared the best-scoring run with the worst. On 9 of 10 tasks, the best run tried more new ideas in its second half. It held up even when the same model ran the same task twice: in 11 of 16 such pairs, the higher-scoring run tried more new ideas late. We can’t say from this that exploring late is what earns the better score, only that the two tend to show up together.
Going back to an old idea paid off less and less. On chip design, an idea’s fourth and later tries made up 36 percent of all experiments. They succeeded about as often as the first three tries, 33 percent against 28, but brought in only 10 percent of the real gain.
Trying new ideas takes time, and the agents were surprisingly quick to give time back. No chip-design run used its full budget. The median run stopped after 19 percent of it, and on every task the best run stopped later than the worst.
Each model has its own research habits
Some agents think at length before each move. Others act in quick bursts and think as they go. Both are reasonable ways to work, and which one an agent uses barely depends on the task.
Change the task, and the habit stays
Of everything each agent wrote, the share that went into actions: commands it ran and code it changed. Each dot is one task.
Six tasks: a math competition, writing, tool calling, medical questions, coding, and science questions. A missing dot means that run’s log lacked what this measure needs.
GPT-5.5 and GPT-5.6 put about a tenth of what they write into actions, whatever the task. Kimi K3 puts in about a third. We saw the same with every habit we measured: knowing which model ran tells you more than knowing which task it was given.^4 One practical consequence is that the prompts, tools and checks built around one model may fit another model badly, and the checks meant to catch an agent’s mistakes should be designed around the habits of the model they watch.
The witness
§5 Grading the Work
[Figure: Ink and watercolor: the same robot walks along the strip of working laid out on a long table, reading it through a big magnifying glass. The part it has read is tinted terracotta; the rest, and a large heap on the floor, is still plain paper. A human finger rests on a line in the read part.]
In the 1660s, witnessing an experiment meant being in the room. With AI researchers the whole record is written down already, and the work left is reading it. Reading 119 of these records turned up things no final score shows, from wins that were only noise to habits that follow a model from task to task. We suspect there is a lot more in the logs nobody has read yet.
There is plenty it doesn’t do yet. It has only seen three benchmarks, and the language model that reads the traces still disagrees with itself now and then (note 5 has the numbers). If you point it at your own agents, we would like to hear what breaks.
In The Goal of Science Is Not to Win, we argued that the process of research deserves as much attention as its result. This is our first try at measuring it.
Cite
@article{shi2026oeb,
title = {Open-Endedness Bench: Measuring Epistemic Process from Agent Records},
author = {Shi, Chengyang and Ji, Xianglin and Huang, Jintao and Wang, Jicheng and He, Yifeng and Liu, Jiachen},
year = {2026}
}
R. W. Wood, “The n-Rays,” Nature 70, 530–531 (1904). ↩
119 runs over 12 tasks: 91 on PostTrainBench (post-training a small language model), 20 on Chip-Bench (hardware design) and 8 on the nanoGPT speedrun. We ran no new agents. The language model that reads the traces was GLM-5.3 for PostTrainBench and Chip-Bench and Grok-4.6 for the speedrun. ↩
A speedrun experiment counts as a real improvement when its loss, at the same number of training steps, falls below the run’s best so far by more than 0.0013, the run-to-run noise of one configuration. On Chip-Bench, an experiment the agent called an improvement beat the run’s best 29% of the time; one it said nothing about, 31%. ↩
For each habit we asked how much of the difference between runs is explained by the model and how much by the task. Across the six habits, the median is 43% for the model and 7% for the task. For thinking versus doing, the habit in the chart, it is 75% and 3%. The chart averages the one or two runs each model made on each of the six PostTrainBench tasks. ↩
When the same language model scored the runs a second time, a score moved by a median of 0.04 on a scale from 0 to 1; with a different language model reading the traces, by 0.08. When a run gives a measure fewer than five cases to work with, the measure reports no score. ↩
