arrow
Return

Deep learning with missing data

delete2026-07-29
delete0
delete
OA
AI
T
Tianyi Ma
T
Tengyao Wang
R
Richard J. Samworth *
DOI:10.1093/jrsssb/qkag114delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In the context of multivariate nonparametric regression with missing covariates, we propose pattern embedded neural networks (PENNs), which can be applied in conjunction with any existing imputation technique. In addition to a neural network trained on the imputed data, PENNs pass the vectors of observation indicators through a second neural network to provide a compact representation. The outputs are then combined in a third neural network to produce final predictions. Our main theoretical result exploits an assumption that the observation patterns can be partitioned into cells on which the Bayes regression function behaves similarly, and belongs to a compositional Hölder class. It provides a finite-sample excess risk bound that holds for an arbitrary missingness mechanism, and in combination with a complementary minimax lower bound, demonstrates that our PENN estimator attains in typical cases the minimax rate of convergence as if the cells of the partition were known in advance, up to a poly-logarithmic factor in the sample size. Numerical experiments on simulated, semi-synthetic, and real data confirm that the PENN estimator consistently improves, often dramatically, on standard neural networks without pattern embedding. Code to reproduce our experiments, as well as a tutorial on how to apply our method, is publicly available.

Journal

J
Journal of the Royal Statistical Society Series B: Statistical Methodology
IF:
0
Papers:
111
Citations:
0

Organization

L
London School of Economics and Political Science
Scholars:
23
Papers: 20
Citations: 0
U
University of Cambridge
Scholars:
986
Papers: 474
Citations: 0
Cited Papers

Cited Papers

Sparse spectral estimation with missing and corrupted measurements
errStat
IF0
err2019-06-26
err0
errOAAI
errAndreas Elsener; Sara van de Geer
errShare
errSave
A Selective Overview of Deep Learning
err2021-05-01
err72
errOAAI
errFan, Jianqing; Ma, Cong; Zhong, Yiqiao
errShare
errSave
mice: Multivariate Imputation by Chained Equations inR
err2011-01-01
err0
errOAAI
errStef van Buuren; Karin Groothuis-Oudshoorn
errShare
errSave
On the consistency of supervised learning with missing values
err2024-09-12
err0
errOAAI
errJulie Josse; Jacob M. Chen; Nicolas Prost; Gaël Varoquaux; Erwan Scornet
errShare
errSave
Highly accurate protein structure prediction with AlphaFold
err2021-07-15
err0
errOAAI
errJohn Jumper; Richard Evans; Alexander Pritzel; Tim Green; Michael Figurnov; Olaf Ronneberger; Kathryn Tunyasuvunakool; Russ Bates; Augustin Žídek; Anna Potapenko; Alex Bridgland; Clemens Meyer; Simon A. A. Kohl; Andrew J. Ballard; Andrew Cowie; Bernardino Romera-Paredes; Stanislav Nikolov; Rishub Jain; Jonas Adler; Trevor Back; Stig Petersen; David Reiman; Ellen Clancy; Michal Zielinski; Martin Steinegger; Michalina Pacholska; Tamas Berghammer; Sebastian Bodenstein; David Silver; Oriol Vinyals; Andrew W. Senior; Koray Kavukcuoglu; Pushmeet Kohli; Demis Hassabis
errShare
errSave
researcher View more