How Hurricane Forecasters Decide When to Trust an AI Model

Two hurricanes were under warning on the morning of Oct. 9, 2026, and the US National Hurricane Center wrote up both in the same forecast cycle. Isaias was over the Gulf and still strengthening, under a life-threatening storm surge warning. Simon was off southwestern Mexico, intensifying fast toward a coast under a hurricane warning. In the discussion for Isaias, the forecaster wrote that the new track sat very near the Google DeepMind guidance, which had been very accurate for the storm. In the discussion for Simon, another forecaster shifted the track away from that same guidance. Neither product says the model was wrong. What separates them is a judgment about guidance, made storm by storm.
The guidance in question is WeatherNext Cyclones, an artificial intelligence model its builders at Google DeepMind describe in Nature, published Aug. 6, 2026. Ferran Alet, Peter Battaglia and colleagues set out to close a gap that has divided forecast models for decades. A global model captures the large-scale flow that steers a storm and so predicts its path well, but it is too coarse to resolve the core where wind speed is decided. A high-resolution regional model has the opposite problem. Besides the Google DeepMind team, the author list carries hurricane center forecasters, among them Wallace Hogsett and John Cangialosi, along with scientists from Colorado State University and the UK Met Office.
The model is already inside the number forecasters start from
Those two sentences matter because the model is not beside the official forecast. It is inside it. The hurricane center's Track and Intensity Models page lists it as GDMN/GDMI, the Google DeepMind Ensemble Mean: the average of 50 ensemble members, trained on decades of reconstructed global weather together with an international archive of observed storms. It is the only AI model in that table that also forecasts wind radii, the storm's size. And it appears in the membership of three of the agency's consensus products, one each for track, intensity and size.
A consensus product is an average of several guidance models. The paper notes that such averages typically beat the models that go into them. Being listed in one is a different status from being available: the agency defines its track consensus as a simple average of at least two named models, and GDMI is first on that list. The paper adds that experimental forecasts were supplied to the hurricane center through the 2025 Atlantic season, and that forecasters there "have already begun incorporating its guidance into their official forecast process."
The numbers belong to the people who built it
The performance claim behind that adoption is the authors' own, and its scope is narrower than it sounds. Evaluated on storms from 2023 to 2025, the team reports that the model's track, intensity and size forecasts "offer an average lead-time advantage of 1 day or more over leading operational models." The models it was measured against are named: the European Centre for Medium-Range Weather Forecasts ensemble, known as ENS, on track, and NOAA's high-resolution regional hurricane model on intensity and size. The hurricane center's official forecast is not one of them, and the paper never claims to beat it.
The sharpest figure is at five days out. There the authors report an average error of 230 kilometers in the model's ensemble-mean track, against 370 kilometers for the ENS ensemble mean on the same storms and the same cycles. Read as lead time rather than distance, that gap is the part worth holding: ENS does not come down to that error until about 3.75 days, so the authors put the newer model at "just over 30 h more advanced warning at that level of accuracy." On intensity they report lower error than the regional model at every lead time they tested.
Every one of those numbers is the developer's measurement of its own model. The paper says so: the work was done under a cooperative research agreement with NOAA, several of the authors have filed a patent on the underlying probabilistic model, and 22 of them are Google employees who hold stock in Alphabet. But no outside group has published its own measurement of the model's performance.
Why the same model gets different weight
Which brings Oct. 9, 2026, back into focus. The Isaias discussion that morning carried no hedge about the guidance at all. The Simon discussion, written for the same cycle, is not a verdict on artificial intelligence either. The forecaster first lists the aids that agreed on the storm approaching the west-central coast of Mexico: the high-resolution global models, NOAA's regional hurricane models, and NOAA's own AI ensemble mean. Then the ones that did not. The discussion reads: "The HCCA and the Google DeepMind ensemble mean have been somewhat inconsistent during the past 2 cycles, windshield-wipering with a track either inland or just offshore beyond the 48-hour period." To hold that down, it says, the forecast blends the various consensus aids and is nudged to the right of both solutions.
Two things in that paragraph are easy to merge and should not be. HCCA is not the AI model. It is the HFIP Corrected Consensus Approach, a corrected average over eleven named models, and the agency's own member list for it does not include GDMI. The complaint is also about consistency between successive cycles on track, not about accuracy: in the same discussion the intensity forecast was hedged high and was "in best agreement with the HFIP-HCCA model." Nothing in either product establishes that any model was wrong. When they were issued, Isaias had not made landfall and Simon had not reached the coast.
The models page states the boundary plainly: "On average, NHC official forecasts have smaller errors than any individual model," and users "should consult the official forecast products issued by the NHC and local National Weather Service Forecast Offices rather than relying solely on forecast model output."
The paper's own argument for why this model belongs in a consensus is not that it is the best one. It is that it is wrong in different places. The authors report that its errors correlate less with the conventional models' than those models' errors do with each other. A guidance product that fails like everything else adds nothing to an average. One that fails differently is worth having in the blend, and worth discounting on the cycle when it cannot hold a track. On the morning of Oct. 9, 2026, both of those things were true of the same model at the same time.
