Insider Brief
- Google DeepMind’s WeatherNext Cyclones AI model outperformed leading operational systems in forecasting tropical cyclone tracks, intensity and wind extent, providing an average lead-time advantage of a day or more.
- At five days, the model recorded an average track error of 230 kilometers, compared with 370 kilometers for ECMWF’s ENS system and 335 kilometers for Google DeepMind’s GenCast.
- The model can generate ensembles of up to 1,000 possible storm scenarios, helping forecasters estimate uncertainty and the probability of damaging winds and rapid intensification.
Google DeepMind researchers have developed an artificial intelligence weather model that they say can predict the path and strength of tropical cyclones more accurately than leading operational systems, potentially giving forecasters a critical day or more to prepare communities for dangerous storms.
The system, called WeatherNext Cyclones, or WN-C, produces forecasts of a storm’s track, maximum wind speed and size for as long as 15 days. In tests covering tropical cyclones from 2023 through 2025, the researchers reported that the model generally outperformed leading conventional and AI forecasting systems.
The results were reported in a peer-reviewed study accepted for publication in Nature. The research was led by scientists at Google DeepMind and included researchers from Google Research, the U.S. National Hurricane Center, Colorado State University and the U.K. Met Office.
At a five-day forecast horizon, WN-C’s average track error was about 230 kilometers, compared with 370 kilometers for the European Centre for Medium-Range Weather Forecasts ensemble system, known as ENS, and 335 kilometers for Google DeepMind’s earlier GenCast AI model. ENS did not reach the same 230-kilometer accuracy until roughly 3.75 days before the storm, translating into a little more than 30 hours of additional warning from WN-C.
The researchers reported similar gains in predicting storm intensity, an area where previous global AI weather models have struggled. At three days, WN-C’s average intensity forecast was 3.75 knots more accurate than the National Oceanic and Atmospheric Administration’s Hurricane Analysis and Forecast System, or HAFS, a high-resolution model designed specifically for tropical cyclones.
The improvements in track and intensity forecasting were comparable to roughly a decade of progress in conventional numerical weather prediction, according to the study.
The findings could have practical implications because hurricane forecasting involves more than predicting where the center of a storm will travel. Emergency managers also need to know how strong a cyclone could become and how far dangerous winds may extend from its center.
“Predicting how dangerous cyclones develop is a longstanding challenge where every hour counts,” the Google DeepMind team writes in a blog post. “Tropical cyclones — also known as hurricanes or typhoons — are among the most destructive weather phenomena on Earth, responsible for more than 700,000 deaths and $1.4 trillion in economic losses globally over the past 50 years. For forecasters, issuing timely, accurate warnings is a constant race against time.”
One Model for Track and Intensity
Tropical cyclone forecasting has long faced a trade-off between scale and detail, according to the scientists.
Global weather models are particularly good at forecasting the large atmospheric patterns that steer hurricanes and typhoons. But because they divide the atmosphere into relatively large grid cells, they can have difficulty representing the small-scale processes around a storm’s core that influence its strength.
Specialized regional models such as HAFS operate at much higher resolution and can better represent those processes. The added detail, however, requires substantial computing resources and limits the area the models can cover.
Previous global AI models faced a similar problem. They could forecast large-scale weather patterns and cyclone tracks accurately but tended to underestimate the strongest winds in tropical cyclones.
WN-C takes a different approach. Instead of learning only from gridded atmospheric data, the researchers jointly trained it on decades of global weather information and a specialized database containing observations from nearly 5,000 tropical cyclones.
That database, the International Best Track Archive for Climate Stewardship, contains information including storm locations, maximum sustained winds, minimum sea-level pressure and the distances that winds of different speeds extend from the storm center.
The model therefore learns large-scale atmospheric behavior while simultaneously learning the characteristics of actual tropical cyclones.
The approach produced a potentially significant result. WN-C uses atmospheric inputs with a resolution of about 0.25 degrees, or roughly 28 kilometers near the equator, yet it still produced competitive intensity forecasts.
The researchers said the finding indicates that extremely high spatial resolution may not be an absolute requirement for accurate hurricane intensity forecasts. Instead, coarse global atmospheric data may contain more information about storm intensity than previously recognized.
Forecasting a Range of Possibilities
WN-C is also an ensemble model, which means that rather than producing a single prediction of what a hurricane will do, it generates multiple plausible versions of the future.
A standard forecast can contain 50 members, with each representing a possible evolution of the global atmosphere and its tropical cyclones. Because the AI model runs relatively quickly, the researchers were also able to generate ensembles containing as many as 1,000 members.
Large ensembles are particularly important in this case because hurricanes are chaotic systems. Small differences in atmospheric conditions can produce substantially different outcomes several days later. Generating many plausible futures allows forecasters to estimate the probability of events rather than relying on one predicted track.
The researchers found that WN-C’s predicted uncertainty was generally well calibrated, meaning that the spread of possible outcomes tended to correspond reasonably well with the model’s actual errors.
The system also improved forecasts of rapid intensification, defined in the study as an increase of at least 30 knots in maximum sustained winds within 24 hours. Rapid intensification remains one of the more difficult problems in hurricane forecasting because a storm can become significantly more dangerous shortly before landfall.
In tests covering the Atlantic and eastern Pacific, WN-C provided a better balance between detecting rapid intensification events and avoiding false alarms than leading individual operational models across much of the range evaluated. Its Critical Success Index, a measure that accounts for correct forecasts, false alarms and missed events, improved from below 0.3 for comparison models to about 0.5.
The system can also estimate the probability that locations will experience winds above 34, 50 or 64 knots. The researchers found that increasing an ensemble from 50 to 1,000 members was particularly useful for estimating rare but potentially costly events several days in advance.
Computational speed makes those larger ensembles practical. According to the study, a WN-C forecast is roughly 10 times faster to generate than one from GenCast and orders of magnitude faster than forecasts from traditional numerical weather prediction systems.
Moving Into Operational Forecasting
The model is already moving beyond retrospective testing, the team reports.
Google DeepMind has publicly released real-time WN-C predictions through its Weather Lab since June 2025. During the 2025 Atlantic hurricane season, experimental forecasts were also supplied to the National Hurricane Center, where forecasters began incorporating the model as additional guidance in their forecasting process, according to the study.
The researchers tested what would happen if WN-C were added to consensus forecasts similar to those used operationally by the National Hurricane Center. Such systems combine several models because averaging forecasts can reduce errors produced by any individual model.
Adding WN-C improved the simulated track consensus by between about 18% and 38%, depending on forecast lead time, with an average improvement of 28%. Adding it to the intensity consensus produced improvements ranging from about 5% to 14%, averaging 6%.
“Our research has already had real-world impact,” the researchers write in the blog post. “During the 2025 hurricane season, our model helped the NHC to make a historic forecast for Hurricane Melissa by predicting the storm’s rapid intensification and landfall in Jamaica. This enabled the NHC to issue an advance warning, giving teams on the ground critical time to prepare. This year, we continue to work together and are now predicting 1,000 possible scenarios for each cyclone to help support forecasters in their decision-making.”
Those results also underscore a limitation of viewing AI as a replacement for conventional weather models. The researchers found that traditional numerical models continued to add useful information, particularly for hurricane intensity. WN-C was valuable partly because its errors were not identical to those made by conventional systems.
The study has other limitations. WN-C depends on high-quality atmospheric analyses produced through existing weather-observation and data-assimilation systems. It therefore does not yet operate directly from raw satellite, aircraft, radar and other observations.
Its local wind estimates also rely on approximate representations of how far winds extend from a cyclone. Actual wind fields can be more irregular than those representations suggest.
And while the model predicts track, intensity and storm structure, other hazards remain outside its current scope. The researchers identified cyclone rainfall, storm surge and precise wind-gust forecasts as areas for future development.
They also said future AI systems could move closer to end-to-end forecasting by directly incorporating raw observations rather than depending on processed atmospheric analyses.
The broader implication could extend beyond hurricanes. The researchers said the same approach — combining large global weather datasets with specialized records of relatively rare extreme events — could potentially be adapted to hazards such as extreme rainfall and localized heat events.
For operational forecasters, however, the immediate significance is more focused. WN-C suggests AI models may be able to combine accurate global storm tracks with useful intensity forecasts without the computing demands traditionally associated with very high-resolution hurricane models.