The Reasoning You Can Read Is Not the Reasoning That Counted

The paper opens with two sentences a small language model wrote while working through a competition geometry problem. First: "Alternatively, maybe I made an error in the problem setup. Let me check again." Later, further down the same trace: "Alternatively, maybe my error is in the coordinate assignment for point A." As text, the two are the same gesture: a model stopping to doubt itself. Measured, they are nothing alike. The first did not change the model's chance of reaching its correct final answer at all. The second caught a sign error in the placement of a single point and raised that chance from about 52% to about 94%.
That gap is what a preprint posted to arXiv on 3 September sets out to measure. Kevin Du of ETH Zürich wrote it with Alexander Hoyle, also at ETH; Laura Ruis of MIT; and Acyr Locatelli of Cohere, and its title says the argument out loud: legibility is not interpretability. The work has not been peer-reviewed; a note from the authors on the posting says it has been accepted at the Conference on Language Modeling.
The distinction is not housekeeping. A good deal of current practice reads step text as though it reported on itself. Models are asked to grade another model's reasoning step by step, to score whether a trace is faithful to what the model actually did, and (the consequential one) to supply the reward that trains the next generation of reasoning models, one step at a time. All of it rests on the words of a step carrying information about the job that step did.
Du and his colleagues propose measuring what a step is worth instead of reading it. They borrow a quantity from reinforcement learning, the trial-and-reward training behind these models: a step's advantage, meaning how much including it improves the chance of a good outcome. Take the model's own trace and cut it on either side of a given step, then rerun the model 50 times from each cut. If the runs that include the step land on the answer more often than the runs that stop just before it, the step earned the difference. A step counts as consequential when it moves the odds by more than a tenth and that shift survives a check against the randomness of the sampling. The models under the microscope are small open ones from the Qwen3 family, working on standard math benchmarks that run from grade-school word problems up to olympiad qualifiers.
Two things emerge before any judging happens. Consequential steps are rare (1.8% of steps in the main test set), so most of what a model writes while reasoning does not measurably change where it ends up. And the gains that come from bigger models, and from thinking mode, do not come from better reasoning along the way. The share of traces that were already on course at the first step and stayed there rose from about a quarter to 61% when thinking mode, the model's long internal draft, was switched on. The improvement is mostly in what the model brings to the problem, not in what it works out while writing.
Next, the authors ran a test: four language models were shown the traces and asked to rate how much each step mattered. Out of the box they did poorly: better than chance as they grew larger, but far below the ceiling the measurement itself allows, since the answer key is estimated from a limited number of reruns and carries noise of its own.
Fine-tuning the same models as step-level critics helped, and it helped lopsidedly. On a standard score for how well a ranked list surfaces rare items, the critics reached about 0.3 on responses that ended in a wrong answer, roughly half of what the measurement allows, and 0.065 to 0.10 on responses that ended in a right one, against a ceiling of much the same height. Making the critic bigger barely moved the second number.
The authors do not oversell that gap. Part of it may be an artifact: a wrong answer often ends with the model simply announcing the wrong result, and a step like that is easy to flag, which flatters the score. But they also make the case that the correct half is the half worth having. Those are the steps where a model finds its way to a solution: exactly what a step-by-step reward model would want to catch, and exactly what the text does not appear to give up.
The reruns behind every measurement have been released as a public dataset, with the code and a viewer for stepping through individual traces, so the labels can be checked rather than taken on trust. The authors close on a problem they do not solve: how to convey that certain reasoning steps influence answers without implying that their text can be read as meaningful.
Sources
- arXivPreprint
