Skip to content
See the World Through ScienceA project of ALLATRA
Source: Peer-reviewed2 sources

Teaching a Speech Decoder That a Wave Is Not a Word

By Oli KotykWriterAI & Technology5 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

Four rows pair a target, the word Hello plus a waving hand, with an example decoded output, naming each mistake as a false negative or false positive for speech or gesture. Box plots of false-positive rates sit below.
The four ways two decoders sharing one electrode grid get it wrong, from missing a wave entirely to reporting speech nobody attempted (panel a of the paper's Fig. 4, above the false-positive rates in panel b).Fig. 4 from Samantha C. Brosler, Jessie R. Liu, Alexander B. Silva, Irina P. Hallinan, Cady M. Kurtz-Miott, Jonah F. B. Dunkel Wilker, Adelyn Tu-Chan, Karunesh Ganguly, Edward F. Chang (2026), "Simultaneous speech and gesture decoding for multimodal communication in paralysis", Nature Neuroscience — CC BY 4.0, cropped · CC-BY-4.0

A decoder that reads attempted speech off the surface of the brain has a blind spot that never mattered while it was the only decoder in the room: it does not know what to do when its user moves. Put a second decoder beside it, one reading attempted hand gestures from the same grid of electrodes, and the two begin answering each other's questions. The gesture model reports a wave while the person is only talking. The speech model prints a phrase while the person is only waving.

How badly that happens, and what it takes to stop it, is the substance of a study published Sept. 14, 2026, in Nature Neuroscience by first author Samantha C. Brosler, senior author Edward F. Chang and colleagues at the University of California, San Francisco. The visible result is a man with ALS driving a personalized full-body avatar on a screen, speaking and gesturing at once. Underneath it sits a less photogenic one: a measurement of how much two decoders sharing one patch of cortex get in each other's way, and a pair of training changes that clears most of it.

Each participant in the trial has a grid of 253 electrodes lying on the surface of the left hemisphere, spanning the strip of cortex that issues movement commands. Movement on that strip is laid out roughly by body part: hand and arm high on the grid, the muscles of speech lower down toward the fold above the ear. That coarse map is what makes a single implant plausible for more than one job, and it held in every participant.

The map is only roughly a map

Roughly is the operative word. Comparing which electrodes fired during attempted speech alone against those active during attempted gestures alone, the team found a set of contacts (concentrated in the precentral gyrus, the front lip of that strip) that responded strongly to both. They scored every electrode for how much it overlapped and mapped the result. In the participant with ALS the overlapping contacts spread across a wide area; in the other, they clustered tightly. Either way, part of the grid was listening to two conversations at once.

The team measured the cost of that overlap in both directions. Decoders whose rest class had only ever seen true rest, the quiet between trials, read the other behavior as an instruction: across the pair of decoders, 30.6% of opposite-behavior trials in the participant with ALS, and 76.0% in the participant paralyzed by a brainstem stroke, produced a spurious word or gesture. Trained only on isolated attempts, the same decoders made the opposite error as well, calling genuine simultaneous attempts rest 14.5% and 34.9% of the time.

Teaching each decoder that the other one is silence

The first fix is close to a pun on the problem. If the gesture decoder keeps hearing speech as a gesture, teach it that speech is silence. The team retrained each model's rest class on a balanced mix of true rest and the other behavior's trials, labeled as rest: speech recordings fed to the gesture model as examples of nothing happening, and gesture recordings fed to the speech model the same way. False alarms on the opposite behavior fell to 0.0% in both participants. That zero is a median, and the intervals around it reach 9.1% and 2.7%. False alarms during true rest stayed as low as they had been.

The second change concerns what the models are trained on: those trained on a mix of isolated and simultaneous recordings held up in both settings, where models trained on either alone did not. With both changes in place, the rate at which a real attempt was mistaken for rest during simultaneous trials fell below 5% in both participants.

Eight-panel scientific figure. Box plots compare gesture and speech decoding accuracy for two participants under different training sets, next to brain maps showing which electrodes each model relied on.
Decoding accuracy for the speech and gesture decoders in two participants, and the electrodes each model leaned on after training on simultaneous attempts. Fig. 3 from Samantha C. Brosler, Jessie R. Liu, Alexander B. Silva, Irina P. Hallinan, Cady M. Kurtz-Miott, Jonah F. B. Dunkel Wilker, Adelyn Tu-Chan, Karunesh Ganguly, Edward F. Chang (2026), "Simultaneous speech and gesture decoding for multimodal communication in paralysis", Nature Neuroscience. CC BY 4.0, resized

Why the mixture helps is partly visible inside the models. When the team compared how much each electrode contributed before and after training on simultaneous attempts, the largest drops landed on exactly the high-overlap contacts, and the largest gains on electrodes that answered to one behavior and not the other. The models learned to lean on the separable signal. The authors are careful with that inference: many overlapping electrodes still contributed, and overlap at a single contact does not mean the two behaviors are indistinguishable in the population.

The decoders themselves are deliberately unremarkable. Each is an ensemble of ten networks trained on different folds of the data, their predictions averaged. A single network runs a short filter across time, then stacked recurrent layers, a standard sequence design. Its inputs are the fast and slow components of the signal from the same electrodes. Running side by side in real time, on a vocabulary of ten phrases and ten gestures attempted after a go-cue, they picked the right phrase 70.0% of the time and the right gesture 66.0%, against 9.1% by chance. Those two figures belong to one person, the man with ALS; the second participant's simultaneous performance was measured offline.

Edward F. Chang and Jessie R. Liu, a co-first author, are named on pending UCSF patent applications covering the neural-decoding approaches described, and Chang cofounded Echo Neurotechnologies.

The system still waits for a go-cue

The boundaries of the result are set by its own design. Every trial follows a countdown and a go-cue, the vocabulary is closed and small, and the authors label the whole system a proof of concept. Three people are enrolled in the trial; the simultaneous task ran with two of them. Generalization to concurrent attempts, by the authors' own assessment, remains imperfect; movements on the same side of the body as the implant decode less accurately; and cutting the lag by decoding continuously rather than trial by trial is named as future work.

What the work hands the field is narrower and more portable than the avatar: a training recipe. The interference the team measured is not peculiar to electrodes resting on the cortical surface. The same problem has appeared in implants that record from inside the tissue, where attempted speech degraded cursor control and the reverse, in models never trained for the two-task case. On this evidence, the limit on an implant that does more than one thing is less about how many electrodes it carries than about whether its decoders have ever been shown the situation they will be used in.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

Teaching a Speech Decoder That a Wave Is Not a Word

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.