A Brain Decoder Was Partly Reading the Clock

The word "the" is short. "Supercalifragilisticexpialidocious" is not. Every listener registers word length without noticing, and that alone is enough to account for most of what one influential brain-reading system appeared to be learning from the brain.
That is the argument of a preprint posted September 30, 2026 by Dulhan Jayalath and Oiwi Parker Jones of the University of Oxford. Their subject is a pipeline published in Nature Communications in November 2025 by Stéphane d'Ascoli, Jean-Rémi King and colleagues, which decodes individual words from magnetoencephalography, a recording of the faint magnetic fields a working brain gives off. That method was a visible step forward, and it has since been extended, benchmarked and reviewed across a line of follow-up work.
Here is how it reads a sentence. A subject listens to an audiobook while a MEG scanner runs. For each word, the system takes a three-second slice of the recording beginning at the moment that word started, and a neural network encodes all of a sentence's slices together before predicting every word at once. Done that way rather than one slice at a time, word classification improved by about half.
Ordinary speech fits several words into three seconds, so a slice cut at one word runs straight through the next few. Neighboring slices therefore hold the same stretch of signal, displaced by the gap between the two word onsets. The Oxford authors write that recovering that gap is a matter of sliding one slice against the other until the shared samples line up, and that the gap is a close proxy for how long the word lasted. A network that picks up on this can narrow its guesses without consulting the brain at all. Shortcut learning is the name for it: a rule that works on the test and does not transfer to the job.

The test that tells a brain signal from a clock
Jayalath and Parker Jones built a control that keeps the timing and discards everything else. They replaced the MEG with a synthetic signal unrelated to anything the subject heard, then cut windows from it at the real word onsets, so neighboring windows overlapped exactly as before. Decoded jointly, that brain-free signal scored 22.0% balanced accuracy, corrected for word frequency, over the fifty most frequent words in the dataset, against 22.3% for real MEG from the same listener. Generate a fresh synthetic signal for each window instead, leaving the overlap with nothing to reveal, and the score drops to 5.8%. A model given the intervals between words as its only input scored 22.9%, a shade better than the real recordings.
Those figures come from a single listener in the LibriBrain100 MEG dataset. The same trained model reproduces the effect on every other subject in that dataset and on two other MEG datasets, though that is the same lab rerunning its own analysis rather than an outside check.
One detail makes the case neatly. In the 2025 paper, the only datasets that leaked no duration were those using reading protocols that gave every word the same time on screen. Those are also the only ones where the original authors found no benefit at all from decoding words together.
How far the problem spreads is a question the authors keep narrow. The construction has been picked up by a line of later brain-to-text work, so the issue travels with it, but how much it matters depends on how informative timing is in each task. They ran their control on one prominent relative of the method, a system that decodes typing from brain activity, and report that timing explains little of its performance.
Removing the shortcut made the brain data count for more
The figure that keeps this from being a demolition is the second rung of the ladder. Real MEG with the windows decoded independently, with no overlap left to exploit, scores 9.5%. The brain-free signal under exactly the same treatment scores 5.8%. The recordings do carry information about which word was heard. It was just far smaller than the headline implied, and it was being drowned out.
So the authors made what they call "a single, simple change": encode each window on its own rather than jointly across the sentence. Their replacement network, SimpleB2T, is roughly a tenth the size of the one it stands in for. Two techniques that had been disappointing then began to work. Averaging the predictions from several separately recorded responses to the same word helped substantially, and so did scoring candidate sentences with an off-the-shelf language model. The paper's explanation is that predictions driven partly by duration are poor material for either: averaging them mostly sharpens the estimate of a word's typical length, and a language model has little to add to a guess that was half stopwatch.
The best result comes with its conditions attached
On the team's own benchmark, SimpleB2T reached a 36.6% word error rate, and the conditions on that number are heavy. Decoding runs over a closed vocabulary of 92 words, on a core set of 100 short care-related sentences. The authors built that clinically motivated communication benchmark themselves, and the result assumes five separately recorded responses to each word, averaged. Given one response per word, the same system scores 65.6%. The sentences are also assembled rather than heard: each position is filled with a different recorded occurrence of that word from held-out sessions, so no listener ever perceived the sentences being reconstructed.
And the subjects were listening, not speaking or trying to speak. That is why the authors stop short of calling this a working brain-computer interface: they write that they do not interpret the method "as yet demonstrating a practical BCI since it operates on perceived and not imagined speech". They also say the benchmark "is not clinically validated". The comparison with surgical implants carries the same hedge: the recipe is "approaching past invasive speech decoding performance, albeit under different conditions", and the invasive study it is measured against used an even smaller vocabulary.
Sources
- arXivPreprint
- Nature Communications
