Skip to content
See the World Through ScienceA project of ALLATRA
Source: Peer-reviewedNature3 sources

A Language Model Rebuilt to Read Every Letter of Every Word

By Wilkens EtienneWriterAI & Technology5 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

Compartments of a wooden printing case filled with loose metal type, with an assembled line of type lying across them.
Each piece of type carries a single character, and a line of text is built up one character at a time. The retrofitted models read text in much the same spirit, byte by byte (illustrative)."Metal movable type" by Willi Heidelbach, via wikimedia, CC-BY-2.5

Strawberry has three r's in it. A large language model will happily tell you what a strawberry is, where it grows and how to turn it into jam, and then count two. The mistake is not a hole in what the model knows. It is a consequence of a decision taken before any training begins: the model is never shown the letters.

Before a model learns anything, its training text is cut into tokens drawn from a fixed list of words and word fragments. The approach is called subword tokenization, and it works well, but it throws away the inside of each chunk. That matters most where the characters are the content: source code, DNA and protein sequences, and languages whose scripts the vocabulary was never sized for.

Writing in Nature on Oct. 7, 2026, Benjamin Minixhofer and colleagues describe a way to rebuild such a model so that it reads the raw bytes of text instead, and they call the procedure byteification. Minixhofer is at the Allen Institute for AI and the University of Cambridge; the senior author, Valentin Hofmann, is at the Allen Institute and LMU Munich.

Models that read bytes are not new, and they have lagged behind the chunk-based kind in practice. The authors' explanation is that nobody was building them the practical way. A byte-level model was trained from scratch and compared against a chunk-based model also trained from scratch, while the chunk-based models kept moving, improving through better training data, better architecture and better post-training. Starting over each time cannot keep pace with that.

Freeze the big model, grow new parts around it

Byteification runs in two stages. In the first, the large transformer at the heart of the model is frozen while a set of new parts is trained around it. Those parts are a small encoder that turns each byte into a representation, a module that decides where to cut the byte stream into patches, a small decoder and a layer that predicts the next byte. The goal of that stage is to reproduce as exactly as possible what the original model would have done. In the second stage everything is unfrozen and trained together, so the model can begin using the finer detail it can now see.

Schematic of the converted model: tokenization and embedding, a local encoder, a boundary predictor, pooling into patches, the global transformer, depooling and a local decoder, with three comparison panels beneath.
Panel a traces one pass of the converted model, from raw bytes through a boundary predictor that groups them into patches and back out to a next byte prediction. Panels b to d set that design beside a chunk reading model and an earlier byte level one. Fig. 1 from Benjamin Minixhofer, Tyler Murray, Tomasz Limisiewicz, Anna Korhonen, Luke Zettlemoyer, Noah A. Smith, Edoardo M. Ponti, Luca Soldaini, Valentin Hofmann (2026), "Retrofitting language models to operate over bytes", Nature — CC BY 4.0, resized

The whole conversion took 49.1 billion tokens and ran for less than one pass through the training data. The authors put that at under 1% of a typical pretraining budget. The phrase is theirs, and it repays reading carefully: the paper names no reference budget to measure it against, and the comparison is in tokens, not in money or energy.

Four models were converted, all of them open-weight and between 1 billion and 8 billion parameters. Olmo 3 7B became Bolmo 7B and OLMo 2 1B became Bolmo 1B; Qwen3 8B Base became Bwen 8B and Llama 3 8B became Blama 8B. No frontier-scale model and no closed commercial model went through the procedure, and the paper claims none.

Where the gains are, and where they stop

On a 40-benchmark suite, the converted models beat every earlier publicly available byte-level model of comparable size. Bolmo 7B scored 16.5 percentage points higher on STEM tasks than BLT 7B, an earlier byte-level model trained from scratch, offered as one example from one category. Against the chunk-based models they were built from, the claim is narrower: the converted versions come close to matching them, and the one place they clearly win is reading individual characters.

The gains are not uniform, and the paper is specific about where they stop. Bwen 8B, the strongest of the four, improved in every category but one: open-ended question answering, where it slightly trailed an earlier byte-level model called TFree-HAT 7B. On code, the converted models produced a working program more often when allowed sixteen attempts and less often on the first. The architecture carries a price of its own as well. Olmo 3 was given the same extra training on the same data without the change of architecture, as a control, and against that baseline, byteification cost a little on multiple-choice questions, open-ended answers and math. On code at the first attempt it cost a lot.

The letter-counting skill, the one that makes the result quotable, did not arrive from the architecture alone. The models were also fed about 75 million tokens of synthetic character puzzles, roughly 0.04% of the training mix: spelling words out, reversing them, swapping and deleting letters, drawn from a word list kept clear of the test set. The authors report that byte-level models do not otherwise pick the skill up on a schedule this short.

Everything here was measured in-house

The statistics behind these comparisons are carefully done, with resampling across the whole benchmark suite and a correction for the number of comparisons made, which is what makes the word significant mean something here. What nobody has done yet is check the models from outside. Every figure in the paper is the authors' own measurement of the authors' own models on a suite they assembled themselves, and the Nature commentary published alongside it was written by two of the paper's own peer reviewers. Set against that, the training code, the training mix, the evaluation data and the peer review file are all public.

None of this is brand new in public. The same authors posted the work as a preprint in December 2025, carrying the same headline numbers; what the Nature version adds is the peer review and the two further conversions. The Allen Institute has released the new checkpoints along with the frozen-transformer checkpoints from the first training stage, which is the part another group would want in order to try the procedure on a model of its own.

One result points at how little a converted model has to give up. The team took the difference in weights between an instruction-tuned version of Olmo 3 and the plain one, added it to Bolmo, and Bolmo's ability to follow instructions rose to the level of the tuned original with no further training. Whatever has been built around a model, in other words, may not have to be built again for its byte-level twin.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

A Language Model Rebuilt to Read Every Letter of Every Word

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.