arrow
Return

Stable Subsampling under Model Misspecification and Covariate Shift

delete2025-11-06
delete0
PRE
AI
J
Jinjing Yang
S
Shaohua Xu
Z
Zebin Yang
A
Aijun Zhang
Y
Yongdao Zhou
DOI:10.1145/3769077delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The presence of covariate shift between training and test datasets, coupled with model misspecification, can lead to instability in regression predictions across diverse datasets. Meanwhile, training complex models with massive data imposes significant computational burden. In this article, we present a novel model-free subsampling algorithm for stable prediction, which employs uniform design and confounder balancing methods. Our subsampling algorithm aims to find the nearest neighbor subsampling points of uniform design with the goal of minimizing global stability loss, thereby reducing the data volume while achieving stable predictions. Theoretic analyses show that the uniform measure minimizes the maximum integrated mean square error (MIMSE) and the global stability loss evaluates the independence among variables in each candidate MIMSE-optimal subsampled sets. Simulation studies conducted on synthetic datasets, as well as applications on real datasets, demonstrate the superiority of our proposed method under model misspecification and covariate shift.

Journal

ACM Transactions on Knowledge Discovery from Data cover
ACM Transactions on Knowledge Discovery from Data
IF:
4.8
Papers:
1.3K
Citations:
4.4K

Organization

No organization information available