All updates

Opinion: We Built the Agents Faster Than We Built the Rulers

A decade in astrophysics teaches you that a measurement without error bars gets torn apart. AI evaluation has not learned that lesson yet.

An illustration of robots streaming off a glowing high-speed assembly line while, below, a crew of workers slowly hand-builds measuring instruments — a ruler, a gauge, a grid, and a plotted chart.
Agent capability has outrun the instruments we use to measure it.

I spent a decade conducting research in theoretical cosmology and astrophysics. One of the first lessons I learned as a wide-eyed physics graduate student was simple: if you present a measurement, it had better come with error bars. Otherwise, you should expect the result — and perhaps you along with it — to be dismantled in front of your peers, mentors, and, somehow, probably your loved ones. This was true even in informal, internal colloquia. More important, you had to be prepared to spend far more effort defending the systematic uncertainties than presenting the measurement itself. For a decade-long argument over what is, in essence, measurement error, look no further than the conflicting estimates of the universe’s expansion rate known as the Hubble tension.

Since moving into AI research, I have watched the field relearn this lesson in real time. We are building AI systems whose complexity rivals even the zoo of active galactic nuclei (AGNs), some of the brightest objects in the universe and the product of a vast collection of physical processes. Yet we are observing these systems with instruments that can feel more primitive than those available in the late 1950s, when astronomers mistook AGNs for ordinary stars. The term quasar, after all, comes from “quasi-stellar object.” The analogy is imperfect, but the lesson holds: our objects of study are evolving faster than our methods for understanding them.

AI research is still only scratching the surface of what constitutes a proper measurement. Comparative model benchmarks are routinely circulated without error bars, despite the obvious variability and stochasticity of these systems, often by design. The problem is not that AI evaluation contains no scientific rigor. It is that the systems we have created are beginning to outclass the laboratories built to evaluate them. Before celebrating another benchmark score, we should ask what, exactly, it measures — and how much confidence it deserves.

The first question is whether we are properly characterizing and decomposing the sources of uncertainty. Is running the same AI evaluation three times enough to characterize uncertainty? Almost certainly not. At best, that estimates within-task variation, and only if the configuration is genuinely fixed, which it rarely is. What about environmental drift introduced by external tools? Are the tasks sampling the intended real-world parent distribution? How much variance comes from the evaluation judge itself? Does a small change to the prompt produce chaotic behavior? A factor held constant in one experiment is not therefore irrelevant. Any of these sources of uncertainty could dominate. Without at least an order-of-magnitude understanding of them, our error bars are largely meaningless.

Even well-estimated uncertainty cannot rescue a poorly specified experiment. An AI system has an extraordinary number of moving parts: the model, prompt, scaffold, tools, permissions, memory, environment state, skills, tasks, and the evaluator itself — ten, by my count. What is being held constant in our experiments? What is the precise configuration? Is the evaluation actually targeting the construct it claims to measure? If we say that we are measuring a model’s ability to conduct research, what is the operational definition of that ability? Claims that component X improved system Y require evidence that X was the only meaningful change and that Y has been disclosed well enough for the result to be reproduced. Token budgets and proprietary models further complicate both comparison and replication.

We must also stop treating the final answer as the entire experiment. Many evaluations of AI capacity in the life sciences still amount to input/output: submit a question to a black box, compare the response with a predetermined label, and hope to infer the system’s inner workings from the outside. The black-hole analogy is apt: observing an output does not reveal the path that produced it. A correct answer to a difficult question is encouraging, but what counts as correct at the frontier of scientific knowledge? Correctness rests on first principles, assumptions, and the problem-solving trajectory itself. Our evaluations should interrogate those elements, not only the final label.

How did we get here? Capital is part of the answer, but reducing the problem to greed is too easy. Release cycles are fast, agent runs are expensive, and a single score attracts attention. So far, the spaghetti thrown at the wall has stuck often enough to reward speed. We should not abandon exploration or rapid building. But if we want AI to participate meaningfully in scientific discovery, we need a concerted commitment to scientific rigor. It is time to untangle the spaghetti before it hardens into standard practice.

I am not going to offer a complete recipe here, at least not yet. Fortunately, the field is beginning to respond. New science-agent benchmarks are providing public harnesses and reproducible tools, using stepwise verification and realistic workflows, and incorporating sound scientific and statistical methods. Rigorous work exists. It is simply not commonplace where it should be required, and it is not yet fully appreciated across industries. That will change under pressure. In the unforgiving arena of the life sciences, smoke and mirrors will not survive contact with reality.