返回
On inference in high-dimensional regression
DOI:10.1093/jrsssb/qkad001.png)
摘要
En 中文
This paper develops an approach to inference in a linear regression model when the number of potential explanatory variables is larger than the sample size. The approach treats each regression coefficient in turn as the interest parameter, the remaining coefficients being nuisance parameters, and seeks an optimal interest-respecting transformation, inducing sparsity on the relevant blocks of the notional Fisher information matrix. The induced sparsity is exploited through a marginal least-squares analysis for each variable, as in a factorial experiment, thereby avoiding penalization. One parameterization of the problem is found to be particularly convenient, both computationally and mathematically. In particular, it permits an analytic solution to the optimal transformation problem, facilitating theoretical analysis and comparison to other work. In contrast to regularized regression, such as the lasso and its extensions, neither adjustment for selection nor rescaling of the explanatory variables is needed, ensuring the physical interpretation of regression coefficients is retained. Recommended usage is within a broader set of inferential statements, so as to reflect uncertainty over the model as well as over the parameters. The considerations involved in extending the work to other regression models are briefly discussed.
Keyword:
confidence sets of models
fixed design
inducement of sparsity
nuisance parameters
parameter orthogonalization
期刊
J
IF:
3.6
论文数:
1.5K
被引数:
3.2W
机构
引用论文
In Defense of the Indefensible: A Very Naive Approach to High-Dimensional Inference捍卫不可辩护: 一种非常幼稚的高维推理方法
STATISTICAL SCIENCE
IF3.4
EXACT POST-SELECTION INFERENCE, WITH APPLICATION TO THE LASSO精确的选择后推理,并应用于套索
ANNALS OF STATISTICS
IF3.7

