A decade in astrophysics teaches you that a measurement without error bars gets torn apart. AI evaluation has not learned that lesson yet.
AI research is still only scratching the surface of what constitutes a proper measurement. On error bars, the extraordinary number of moving parts, and why the final answer is not the entire experiment.
Scientific language carries strong priors about what is plausible. When semantic plausibility and formal consistency disagree, which one governs the model's reasoning?
A contradiction can be hard because the formal system is hard to solve, or because it is not obvious which formal system the evidence should become. A small experiment separates the two steps.
Adapting a model to an expert community teaches it where that community agrees. What happens to the disagreement within it?
Fine-tuning can reproduce expert answers without reproducing the judgment that produced them. What may be lost is not knowledge, but the model's tendency to consider an alternative on its own.
AI has advanced fastest on tasks that are easier to verify than to solve. What is the analogue of verification in science?
Scientific hypotheses are not established by proof but shaped by evidence. As AI makes hypothesis generation cheap, the scarce resource shifts from ideas to evidence — and knowledge representation has to shift from facts to belief states.
For scientific agents, how sensitive are conclusions to perturbations of the organization of knowledge itself?
Output stability asks whether an agent reaches the same answer. Path stability asks whether it gets there the same way. The trajectory may reveal more about a knowledge organization than the final answer ever could.