arrow
Return

A PARTIALLY LINEAR FRAMEWORK FOR MASSIVE HETEROGENEOUS DATA

delete2016-08-01
delete118
delete
OA
AI
赵天琦 (Tianqi Zhao) *
G
Guang Cheng
H
Han Liu
DOI:10.1214/15-AOS1410delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
We consider a partially linear framework for modeling massive heterogeneous data. The major goal is to extract common features across all subpopulations while exploring heterogeneity of each subpopulation. In particular, we propose an aggregation type estimator for the commonality parameter that possesses the (nonasymptotic) minimax optimal bound and asymptotic distribution as if there were no heterogeneity. This oracle result holds when the number of subpopulations does not grow too fast. A plug-in estimator for the heterogeneity parameter is further constructed, and shown to possess the asymptotic distribution as if the commonality information were available. We also test the heterogeneity among a large number of subpopulations. All the above results require to regularize each subestimation as though it had the entire sample. Our general theory applies to the divide-and-conquer approach that is often used to deal with massive homogeneous data. A technical by-product of this paper is statistical inferences for general kernel ridge regression. Thorough numerical results are also provided to back up our theory.
Keywords:
Divide-and-conquer method
heterogeneous data
kernel ridge regression
massive data
partially linear model
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Annals of Statistics cover
Annals of Statistics
IF:
3.7
Papers:
2.8K
Citations:
2.9W

Organization

P
Princeton University
Scholars:
2.1W
Papers: 2.3W
Citations: 5.1W
Purdue University System cover
Purdue University System
Scholars:
3.9W
Papers: 3.6W
Citations: 66