返回
Cross-prediction-powered inference
DOI:10.1073/pnas.2322083121.png)
摘要
En 中文
While reliable data-driven decision-making hinges on high-quality labeled data, the and expensive scientific measurements. Machine learning is becoming an appealing from satellite imagery are used to supplement accurate survey data, and so on. Since the validity of downstream inferences. We introduce cross-prediction: a method for valid inference powered by machine learning. With a small labeled dataset and a large unlabeled dataset, cross-prediction imputes the missing labels via machine learning and applies a form of debiasing to remedy the prediction inaccuracies. The resulting inferences achieve the desired error probability and are more powerful than those that only leverage the labeled data. Closely related is the recent proposal of Jordan, T. Zrnic, Science 382, 669-674 (2023)], which assumes that a good pretrained model is already available. We show that cross-prediction is consistently more powerful than an adaptation of prediction-powered inference in which a fraction of the labeled data is split off and used to train the model. Finally, we observe that cross-prediction gives more stable conclusions than its competitors; its CIs typically have significantly lower variability.
Keyword:
statistical inference
CIs
machine learning
prediction
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
P
IF:
9.1
论文数:
10.8W
被引数:
73.5W

