arrow
Return

How to deal w___ missing input data

delete2025-11-13
delete0
delete
OA
AI
M
Martin Gauch *
F
Frederik Kratzert
D
Daniel Klotz
G
Grey Nearing
D
Déborah Cohen
O
Oren Gilon
DOI:10.5194/hess-29-6221-2025delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep learning hydrologic models have made their way from research to applications. More and more national hydrometeorological agencies, hydro power operators, and engineering consulting companies are building Long Short-Term Memory (LSTM) models for operational use cases. All of these efforts come across similar sets of challenges - challenges that are different from those in controlled scientific studies. In this paper, we tackle one of these issues: how to deal with missing input data? Operational systems depend on the real-time availability of various data products - most notably, meteorological forcings. The more external dependencies a model has, however, the more likely it is to experience an outage in one of them. We introduce and compare three different solutions that can generate predictions even when some of the meteorological input data do not arrive in time, or not arrive at all: First, input replacing, which imputes missing values with a fixed number; second, masked mean, which averages embeddings of the forcings that are available at a given time step; third, attention, a generalization of the masked mean mechanism that dynamically weights the embeddings. We compare the approaches in different missing data scenarios and find that, by a small margin, the masked mean approach tends to perform best.
Keywords:
DATA SET
CAMELS

Journal

Hydrology and Earth System Sciences cover
Hydrology and Earth System Sciences
IF:
5.8
Papers:
6.0K
Citations:
2.8W

Organization

A
alphabet inc.
Scholars:
1.1K
Papers: 663
Citations: 0
G
Google Incorporated
Scholars:
3.5K
Papers: 1.8K
Citations: 8