Skip to content
See the World Through Science
Source: PreprintarXiv2 sources

AI Weather Models Can Predict the Past, and That Is a Clue to Why They Work

By Oli KotykWriterAI & Technology6 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

A butterfly-shaped tangle of fine orange curves on a black background: the two lobes of the Lorenz attractor.
The Lorenz attractor, drawn from the toy weather equations Edward Lorenz used in 1963. Its butterfly shape gave the butterfly effect its name. Illustrative diagram, not a figure from the study."το χάος : Chaos (Lorenz attractor)" by dullhunk, via flickr, CC-BY-2.0 · CC-BY-2.0

Ask a modern weather model what Wednesday will look like and it will tell you, accurately enough that leading forecasting centers have begun running these systems operationally. Ask it what Monday looked like and, by everything atmospheric physics has to say, you should get nonsense. The atmosphere destroys information as it goes, grinding organized motion into heat through friction, turbulence and cloud. Run that in reverse and the damping becomes explosive growth. Do it with the equations and the numbers blow up immediately.

Neural networks do not appear to mind. In a preprint posted to arXiv on August 26, 2026 Pedram Hassanzadeh of the University of Chicago and colleagues at Nanjing University and New York University report that AI models trained to step the atmosphere backward reconstruct the past skillfully: out to 6.5 days on reanalysis data, against 9.1 days going forward for the same architecture. The work has not been through peer review.

These are separate networks, not the forecaster played in reverse: same architecture, same data, same loss function, with only the input and output pairs swapped. Hand one today's atmosphere and it gives back yesterday's.

The gap between the two numbers is itself one of the paper's results. Backcasts are, in the authors' words, systematically less accurate than forecasts, and the asymmetry holds across variables, levels and skill measures. It grows larger still in a self-contained climate model, where no observations nudge the state back toward reality. The past is harder to predict than the future, just nowhere near as hard as it should be.

Skillful backcasting, the preprint says, "appears to violate the second law of thermodynamics." The hedge is the authors' own and it is doing real work. Nothing here breaks physics: a neural network never integrates the atmosphere's equations; it learns a statistical map from one state to the next, so the constraint that wrecks a numerical model run backward has no grip on it.

The butterfly that never turns up

A second oddity was already on the record. In a truly multi-scale chaotic system, a disturbance too small to measure grows quickly and feeds upward, scale by scale, until it corrupts the largest patterns in the flow. That is the butterfly effect, and it is what sets the two-week ceiling on weather prediction. Physics-based models reproduce it; AI models do not. Hassanzadeh's team finds the same blank in its backcasting models, where theory says growth should be more violent still.

That half does not rest on a single preprint. Tobias Selz and George Craig, at the Karlsruhe Institute of Technology and LMU Munich, scored several state-of-the-art AI forecast systems against six characteristics of the butterfly effect and published the result in June in the peer-reviewed Journal of Geophysical Research: Machine Learning and Computation. The models failed, and the two concluded it "seems likely" the failure "results from limitations in the analysis data used for training," since size, design and architecture "turned out to be largely irrelevant." Their paper corroborates the missing butterfly and its cause. It says nothing about backcasting.

One cause, argued in a toy atmosphere

The suspect is the training data, and specifically how smooth it is. ERA5, the reanalysis record nearly every AI weather model learns from, is itself the output of a numerical model on a grid near 31 km, and the AI is trained to jump six or 24 hours at a stride. The motions that carry error upward are not in there to be learned.

To test whether that absence explains anything, the team moved to a system where they could switch it on and off: a toy atmosphere of eight large, slow variables coupled to hundreds of small, fast ones. They trained a dozen networks on it, all sharing one architecture and one loss function, varying only what the training data contained. Fed the slow variables alone at long time steps, the networks behaved just like the weather models: good forecasts, skillful backcasts, no butterfly. Fed everything the numerical solution holds, the same networks turned physics-like. Backcasts blew up, butterfly-like error growth appeared, and forecast skill got worse by a factor of about seven.

On that evidence the authors propose what they explicitly call three hypotheses, tied to one cause: that the coarse-graining of training data is behind the models' unexpected forecast skill, their missing butterfly effect and their ability to predict the past. The claim is strongest where they could manipulate the cause by hand, in the toy system. For the real atmosphere it stays an inference, drawn from a hierarchy of increasingly idealized models.

The closest thing to a real-world test is the Pangu-Weather family: four official networks trained on the same ERA5 data at time steps of 24, 6, 3 and 1 hour. Run continuously, which had not been done for the shortest steps, the four line up in a row. As the step shrinks, ensemble spread stops growing at a lazy, amplitude-independent rate and starts growing fast from the small scales up, the way a chaotic fluid does, while the skillful forecast range falls from 9.3 days to 2.0.

The authors are careful about what they are and are not claiming. "We emphasize that we are not claiming that this is the (real) butterfly effect," they write. ERA5 is heavily smoothed even at its native resolution, and the one-hour model produces more small-scale variance than ERA5 itself holds, which the team flags as a warning sign. Butterfly-like is the phrase they use throughout.

The trade-off waiting at the long end

Three days out, none of this is a problem. The explanation the paper offers for AI's accuracy is almost flattering: trained on data that never held the small, fast scales, the models learned how those scales push the large ones around without inheriting the runaway error growth that comes with resolving them.

The bill arrives further out. Probability forecasts weeks to months ahead depend on ensemble spread being right, not merely present, and the climate emulators now being built on kilometer-scale simulations learn from output that is coarsened first, to grids around 100 km and daily averages. That is the same smoothing this paper says strips the physics out. The authors pose the question rather than settle it: whether we want AI models to behave like physics at all, if behaving like physics costs the skill that made them useful.

Backcasting itself comes out of the work as an instrument. A model that runs backward is a way to interrogate what a network actually absorbed, and it may prove useful for hunting rare extremes and their precursors, a job now done with limited linear tools. The team's own first attempt at a single model trained in both directions has produced no improvement so far, which they say plainly.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

AI Weather Models Can Predict the Past, and That Is a Clue to Why They Work

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.