返回
Consistent Sparse Deep Learning: Theory and Computation
DOI:10.1080/01621459.2021.1895175.png)
摘要
En 中文
Deep learning has been the engine powering many successes of data science. However, the deep neural network (DNN), as the basic model of deep learning, is often excessively over-parameterized, causing many difficulties in training, prediction and interpretation. We propose a frequentist-like method for learning sparse DNNs and justify its consistency under the Bayesian framework: the proposed method could learn a sparse DNN with at most O(n/ log(n)) connections and nice theoretical guarantees such as posterior consistency, variable selection consistency and asymptotically optimal generalization bounds. In particular, we establish posterior consistency for the sparseDNNwith amixture Gaussian prior, showthat the structure of the sparse DNN can be consistently determined using a Laplace approximation-basedmarginal posterior inclusion probability approach, and use Bayesian evidence to elicit sparse DNNs learned by an optimization method such as stochastic gradient descent in multiple runs with different initializations. The proposed method is computationally more efficient than standard Bayesian methods for large-scale sparse DNNs. The numerical results indicate that the proposed method can perform very well for large-scale network compression and high-dimensional nonlinear variable selection, both advancing interpretable machine learning.
Keyword:
Bayesian evidence
Laplace approximation
Network compression
Nonlinear feature selection
Posterior consistency
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
J
IF:
3
论文数:
5.2K
被引数:
4.8W
机构
引用论文
BAYESIAN ESTIMATION OF SPARSE SIGNALS WITH A CONTINUOUS SPIKE-AND-SLAB PRIOR
ANNALS OF STATISTICS
IF3.7
Bayesian variable selection for high dimensional generalized linear models: Convergence rates of the fitted densities
ANNALS OF STATISTICS
IF3.7
Optimal approximation of piecewise smooth functions using deep ReLU neural networks
NEURAL NETWORKS
IF6.3

