返回
Optimal Subsampling via Predictive Inference
DOI:10.1080/01621459.2023.2282644.png)
摘要
En 中文
In the big data era, subsampling or sub-data selection techniques are often adopted to extract a fraction of informative individuals from the massive data. Existing subsampling algorithms focus mainly on obtaining a representative subset to achieve the best estimation accuracy under a given class of models. In this article, we consider a semi-supervised setting wherein a small or moderate sized labeled data is available in addition to a much larger sized unlabeled data. The goal is to sample from the unlabeled data with a given budget to obtain informative individuals that are characterized by their unobserved responses. We propose an optimal subsampling procedure that is able to maximize the diversity of the selected subsample and control the false selection rate (FSR) simultaneously, allowing us to explore reliable information as much as possible. The key ingredients of our method are the use of predictive inference for quantifying the uncertainty of response predictions and a reformulation of the objective into a constrained optimization problem. We show that the proposed method is asymptotically optimal in the sense that the diversity of the subsample converges to its oracle counterpart with FSR control. Numerical simulations and a real-data example validate the superior performance of the proposed strategy. Supplementary materials for this article are available online.
Keyword:
Convex quadratic programming
Cross-validation
Density estimation
False selection rate
Local false discovery rate
Uniform convergence
期刊
J
IF:
3
论文数:
5.2K
被引数:
4.8W
机构
引用论文
Additional quantum numbers for two-electron states in solids. Application to topological superconductor UPt3固体中两电子态的附加量子数。在拓扑超导体UPt3中的应用
PECVD-grown carbon nanotubes on silicon substrates with a nickel-seeded tip-growth structure具有镍种子尖端生长结构的硅衬底上的PECVD生长的碳纳米管

