Return
Distance Learning for Analog Methods
DOI:10.1175/MWR-D-24-0204.1.png)
Abstract
En 中文
Analogs are similar states of a system, occurring at remote times within independent numerical simulations or previous observations. This concept has emerged in atmospheric sciences and was further used in ocean sciences for forecasting, downscaling, upscaling, and extreme event attribution. The distance used to find and rate analogs is a key fea-ture of analog methods. Most studies are based on the Euclidean distance or other predefined metrics. In this investigation, we adapt distance learning algorithms originally designed for classification and regression to statistical forecasting objec-tives, using the continuous-ranked probability score as a loss function. Our algorithm allows to jointly optimize three key hyperparameters of analog methods: the feature space, distance, and the number of analogs used. In particular, this algo-rithm allows to reduce the feature space dimension while preserving high performances, a key requirement for small-sized datasets. We test our algorithm on an idealized chaotic system and a tropical cyclone dataset. These experiments suggest that the optimal distance depends on the forecast horizon and the number of available data and that our algorithm allows for reasonable performances of analog ensemble methods even for small-sized datasets. Our algorithm runs faster than ex-isting grid-search-like analog distance optimization algorithms. This allows to test and optimize a wider class of distances, for instance, weighting a large number of predictive variables. Our approach is not limited to forecasting and can assist in the search for optimal hyperparameters of any analog method, enhancing exploration possibilities and improving overall performances.
Keywords:
Optimization
Machine learning
Ensembles
Statistical forecasting
Downscaling
Hurricanes/typhoons
Journal
IF:
3
Papers:
103
Citations:
2.9W

