返回
Stable Prediction With Leveraging Seed Variable
DOI:10.1109/TKDE.2022.3169333.png)
摘要
En 中文
In this paper, we focus on the problem of stable prediction across unknown test data, where the test distribution might be different from the training one and is always agnostic when model training. In such a case, previous machine learning methods might exploit subtly spurious correlations induced by non-causal variables in training data for prediction. Those spurious correlations can vary across datasets, leading to instability of prediction across unknown test data. To address this problem, we propose an algorithm based on conditional independence tests to screen out non-causal features and reduce spurious correlations by leveraging a seed variable. We show, both theoretically and with empirical experiments, that our algorithm can precisely screen out the isolated non-causal variables, which have no causal relationship with other variables, and remove the spurious correlations induced by them, increasing the stability of prediction across unknown test data. Extensive experiments on both synthetic and real-world datasets demonstrate that our algorithm outperforms state-of-the-art methods for stable prediction across unknown test data.
Keyword:
Conditional independence
separation
seed variables
stable prediction
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
引用论文
Stable learning establishes some common ground between causal inference and machine learning稳定学习在因果推理和机器学习之间建立了一些共同点

