Skip to content
See the World Through ScienceA project of ALLATRA
Source: PreprintEGUsphere2 sources

A North Sea Flood Forecast Gets Better When Most of the Model Runs Are Discarded

By Anna KotlyarWriterNatural Disasters4 min read

Republish this story

Our work is licensed under Creative Commons BY-NC 4.0. You may republish this piece for free — with credit to ALLATRA Media and a link to the original, unedited beyond length trims, and not for commercial use.

Read the full license

Floodwater covers a fenced riverside promenade in the Port of Hamburg during a storm surge, with container cranes, a ferry and sunbeams breaking through heavy cloud behind.
A storm surge floods the riverside promenade in the Port of Hamburg. Seasonal forecasts try to signal months ahead whether a coming season is likely to bring surges like this one; the photograph shows an earlier event, not a case from the study."Hafen Hamburg bei Sturmflut" by greenoid, via flickr, CC-BY-SA-2.0 · CC-BY-SA-2.0

A seasonal climate forecast is not one prediction. It is a crowd of them: the same model run many times over, each copy nudged from a slightly different starting point, each drifting into its own version of the coming months. The conventional way to read that crowd is to average it, and for storm surge on the North Sea coast, the average says almost nothing.

That is the problem Anna K. Miesner and Leonard F. Borchert of the University of Hamburg's Earth and Society Research Hub, working with Daniel Krieger of the Max Planck Institute for Meteorology, set out to get around. Their paper went online on Sept. 2 as an EGUsphere preprint, open for public comment and under review at the journal Natural Hazards and Earth System Sciences. Nobody outside the three of them has checked it yet.

The average fails not because the individual runs are worthless, but because they disagree about one specific thing: which large-scale weather patterns will dominate the season. Surge on this coast is driven by wind and low pressure, so a run that puts the season's circulation in roughly the right regime will get the surge roughly right, and a run that does not will not. Average the two together and both are erased.

So the team stopped averaging. Working with the high-resolution version of the Max Planck Institute Earth System Model's forecast ensemble, they kept only the members whose circulation resembled the weather patterns known to accompany North Sea surges, then scored what was left.

Skill of this kind is measured as an anomaly correlation: how closely a forecast's year-to-year ups and downs track the ones that actually occurred, on a scale where 1 is perfect and 0 is no better than guessing. Given what the authors call perfect knowledge of weather patterns, told in advance which patterns the season really produced, the selected sub-ensemble reaches 0.78 for seasonal surge height and 0.64 for the number of surge events. That is an upper bound rather than a forecast, the score the method would earn if the hardest part of the job were already done for it.

When the selection has to work from the weather patterns the model itself predicts (the only version anyone could actually run), the two scores drop to 0.27 and 0.31. They are real, and they are weak. Earlier seasonal forecasts of winter storminess for the German Bight, at the southeastern corner of the North Sea, produced correlations in that same range and were reported as not statistically significant.

Against the plain ensemble mean, the subselection is an improvement, the authors report; the full ensemble on its own shows limited skill at these lead times.

The distance between the two pairs of numbers is what the paper is actually about. The authors write that the gap "highlights the potential to improve seasonal storm surge forecasts through better prediction of surge-relevant atmospheric regimes." The place to push, in other words, is the atmosphere. Push the reading one step further and the gap also implies that the second link in the chain is in reasonable shape: hand the method the right patterns and it turns them into surge estimates well. That step is an inference from the numbers, not a claim the paper makes.

The selection trick itself is not new. In 2024, Krieger and colleagues published a peer-reviewed version of the same idea in Natural Hazards and Earth System Sciences, choosing members of the same Max Planck ensemble to forecast German Bight storm activity, the winds rather than the water they push. What is new here is the move to surge itself, both its seasonal height and the count of surge events, and the choice to run the idealized case beside the realistic one so that the distance between them can be read as a diagnosis.

What the paper offers, in its own summary, is "a pathway towards earlier and more reliable coastal hazard predictions," not a seasonal outlook a coastal authority could plan around this year. Whether the pathway survives contact with reviewers is now a public question: discussion on the preprint stays open until Oct. 14, and referees' objections, when they come, will appear beneath the paper for anyone to read.

Sources

Spot an error?

Spot an error?

Report an error

Spotted a mistake on this page? Tell us what's wrong and our editors will take a look.

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We correct mistakes openly. Select any text to flag it. Fixes are logged under our Corrections Policy.

Report an error

Reporting on

A North Sea Flood Forecast Gets Better When Most of the Model Runs Are Discarded

What kind of problem?

Only if you'd like us to be able to follow up. We won't use it for anything else.

We read every report. Corrections are logged publicly.