arrow
返回

Bayesian interpolation with deep linear networks

delete2023-05-30
delete10
delete
OA
AI
B
Boris Hanin *
A
Alexander Zlokapa
DOI:10.1073/pnas.2301345120delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Characterizing how neural network depth, width, and dataset size jointly impact model quality is a central problem in deep learning theory. We give here a complete solution in the special case of linear networks with output dimension one trained using zero noise Bayesian inference with Gaussian weight priors and mean squared error as a negative log-likelihood. For any training dataset, network depth, and hidden layer widths, we find nonasymptotic expressions for the predictive posterior and Bayesian model evidence in terms of Meijer-G functions, a class of meromorphic special functions of a single complex variable. Through asymptotic expansions of these Meijer-G functions, a rich new picture of the joint role of depth, width, and dataset size emerges. We show that linear networks make provably optimal predictions at infinite depth: the posterior of infinitely deep linear networks with data-agnostic priors is the same as that of shallow networks with evidence-maximizing data-dependent priors. This yields a principled reason to prefer deeper networks when priors are forced to be data -agnostic. Moreover, we show that with data-agnostic priors, Bayesian model evidence in wide linear networks is maximized at infinite depth, elucidating the salutary role of increased depth for model selection. Underpinning our results is an emergent notion of effective depth, given by the number of hidden layers times the number of data points divided by the network width; this determines the structure of the posterior in the large-data limit.
Keyword:
deep learning
Bayesian inference
neural networks
linear networks
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

P
Proceedings of the National Academy of Sciences of the United States of America
IF:
9.1
论文数:
10.8W
被引数:
73.5W

机构

P
Princeton University
学者数:
2.1W
论文数: 2.3W
被引数: 5.1W
引用论文

引用论文

Modeling maximum daily temperature using a varying coefficient regression model
err2014-04-10
err0
PREAI
errHan Li; Xinwei Deng; Dong‐Yun Kim; Eric P. Smith
err分享
err收藏
err分享
err收藏
Effect of Water and Chemical Stresses on the Silver Coated Polyamide Yarns
err2019-12-26
err0
PREAI
errEzgi Ismar; Shahood uz Zaman; Xuyuan Tao; Cédric Cochrane; Vladan Koncar
err分享
err收藏
Adhesion molecules and their ligands in chronic rejection of human renal allografts
err1997-02-01
err0
PREAI
errE. von Willebrand; V. Jurcic; H. Isoniemi; P. Häyry; T. Paavonen; L. Krogerus
err分享
err收藏
Micropropagation, seed propagation and germplasm bank of Mandevilla velutina (Mart.) Woodson
err2007-06-01
err0
errOAAI
errRonaldo Biondo; Ana Valéria Souza; Bianca Waléria Bertoni; Andreimar Martins Soares; Suzelei Castro França; Ana Maria Soares Pereira
err分享
err收藏
High-dimensional dynamics of generalization error in neural networks
err2020-12-01
err150
errOAAI
errAdvani, Madhu S.; Saxe, Andrew M.; Sompolinsky, Haim
err分享
err收藏
err分享
err收藏
学者 查看更多内容