Skip to content
See the World Through ScienceA project of ALLATRA
Source: PreprintarXiv2 sources

Bigger AI Models Coped Better With Simulated Hardware Faults

By Wilkens EtienneWriterAI & Technology5 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

A microscope photograph of a chip die showing a dense grid of identical logic cells in blue and purple.
Rows of identical logic cells on a chip die, magnified under a microscope (illustrative of the hardware class). The faults in the study were simulated in software on ordinary GPUs, and no faulty chip was built."Plessey Semiconductor (Ferranti) ULA 2C200J Die Photograph" by Revaldinho, via wikimedia, CC-BY-4.0

The chips that run today's large AI models are kept deliberately overcautious. Hold a logic gate at a high enough voltage and it will almost never flip the wrong way; the catch is that the energy needed to switch that gate rises with the square of the voltage. Reliability is bought with electricity. The brain does not make that bargain. It computes with neurons that misfire routinely, and it does the work on a tiny fraction of the power any artificial network of comparable reach consumes.

On Oct. 7, 2026, Trevor McCourt, Ila R. Fiete and Isaac L. Chuang posted a preprint, not yet peer reviewed, asking whether a language model can be taught to live with that kind of unreliability. McCourt and Chuang work in MIT's department of electrical engineering and computer science, Fiete in brain and cognitive sciences, and their answer, drawn from 40,000 GPU-hours of training runs, is that resistance to hardware errors does not fall away as models grow. It rises.

Nothing in that experiment ran on a defective chip. The faults were emulated in software on ordinary GPUs: inside the attention and feed-forward layers of Llama2-style models, random blocks of four numbers were zeroed out of each matrix multiplication, with a fresh set drawn on every pass and the error rate dialed from zero up to one block in five. It is a stand-in for a processor that drops a connection whenever it catches itself making an arithmetic mistake. Models trained under those conditions are what the paper calls fault-hardened. The ordinary kind, trained in a clean environment, is what it calls fault-blind.

Comparing the two needs a single number, and the one the team settled on is how large a conventional fault-free model you would have to build to match a given fault-hardened one. That ratio is the model's useful capacity: the share of its weights doing work rather than absorbing errors. Thousands of models were trained to map it out on a 350-billion-token slice of FineWeb, a corpus of web text, at sizes running up to 930 million weights.

The curve falls before it climbs

Useful capacity does not simply rise. It falls first as models get bigger, exactly as a pessimist would predict, then turns around and recovers past a critical size the paper labels N*. Capturing that shape meant adding a term to the standard scaling law. A scaling law here is the empirical rule that a model's error drops as a steady power of its size and of how much text it has read, and the usual form has no way to bend the way this data bends; fits without the extra term miss badly.

The explanation on offer starts from biology. Nature protects information in two ways. The crude way is repetition, keeping many copies of a value, and redundancy of that kind eats an ever larger share of a system as it grows. The powerful way is a distributed code, like the one the brain's grid cells use to hold an animal's position, where the overhead stays a fixed fraction at any size. A model that could only manage repetition would watch its useful fraction shrink toward zero. These did the opposite.

Why a model would find such a thing without being told to is less mysterious than it sounds. Learning, like evolution, drifts toward arrangements that survive damage, because a system that can absorb a mistake can afford to try more things.

The authors are careful about what follows. The recovery past N* strongly suggests the models are learning good codes rather than copying themselves; whether a trained model is formally fault-tolerant is a conjecture in the paper, not a result. The fitted laws suggest the overhead of that learned code stays finite however large a model gets, and finite is not the same as small. The runs behind all of it stopped at about a billion weights.

The chip in this story has not been built

The fault model was not picked arbitrarily. It imitates a design McCourt and his colleagues call a Low Energy Neural Network Accelerator: a chip that would run its arithmetic below the voltage needed for reliable operation, catch its own errors, and convert each one into a dropped connection. The trick is that it only detects. Every number going into a multiplication carries a small extra remainder alongside it, the same principle the grid cells use, and a cheap check tests whether the running total is still consistent. Repairing what the check finds is left to the model, because in this arithmetic spotting an error is far cheaper than fixing one.

That accelerator is described in a companion manuscript by the same group, listed in the references as submitted for publication. There is no public version of it, so the hardware half of the energy argument is the half nobody outside the group can read.

What it would take to know

The distance between a billion weights and a commercial model is where the claim is most exposed. Repeating the experiment at 100 billion weights, the scale that would settle it, would cost on the order of 100 million GPU-hours by the authors' estimate, comparable to the entire training budget of a frontier model: a cost they call intractable within academia. They suggest cheaper proxies first, on simple datasets where the far end of the curve is actually reachable. The contrast with fault-blind models, meanwhile, is already stark: past a certain error rate, extra training data made those models worse rather than better when faults were injected as they answered.

Training a network around hardware errors is not new. It has been done at small scale for low-voltage chips, and built into particular billion-weight models at a single size. What is new here is the direction of the trend across sizes. The nearest theoretical precedent, a 2024 proof in Physical Review E that fault-tolerant networks can be built from biological error-correcting codes, came largely from the same MIT group, with Fiete and Chuang among its authors.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

Bigger AI Models Coped Better With Simulated Hardware Faults

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.