返回
Markov Subsampling Based on Huber Criterion
DOI:10.1109/TNNLS.2022.3189069.png)
摘要
En 中文
Subsampling is an important technique to tackle the computational challenges brought by big data. Many subsampling procedures fall within the framework of importance sampling, which assigns high sampling probabilities to the samples appearing to have big impacts. When the noise level is high, those sampling procedures tend to pick many outliers and thus often do not perform satisfactorily in practice. To tackle this issue, we design a new Markov subsampling strategy based on Huber criterion (HMS) to construct an informative subset from the noisy full data; the constructed subset then serves as refined working data for efficient processing. HMS is built upon a Metropolis-Hasting procedure, where the inclusion probability of each sampling unit is determined using the Huber criterion to prevent over scoring the outliers. Under mild conditions, we show that the estimator based on the subsamples selected by HMS is statistically consistent with a sub-Gaussian deviation bound. The promising performance of HMS is demonstrated by extensive studies on large-scale simulations and real data examples.
Keyword:
Markov processes
Estimation
Task analysis
Noise measurement
Convergence
Data models
TV
Markov chain
regression
robust inference
subsampling
期刊
IF:
8.9
论文数:
7.6K
被引数:
7.2W
机构
引用论文
STATISTICAL CONSISTENCY AND ASYMPTOTIC NORMALITY FOR HIGH-DIMENSIONAL ROBUST M-ESTIMATORS高维鲁棒M-估计的统计一致性和渐近正态性
ANNALS OF STATISTICS
IF3.7
A general bahadur representation of M-estimators and its application to linear regression with nonstochastic designs
ANNALS OF STATISTICS
IF3.7

