返回
Deep Neural Network Training With Distributed K-FAC
DOI:10.1109/TPDS.2022.3161187.png)
摘要
En 中文
Scaling deep neural network training to more processors and larger batch sizes is key to reducing end-to-end training time; yet, maintaining comparable convergence and hardware utilization at larger scales is challenging. Increases in training scales have enabled natural gradient optimization methods as a reasonable alternative to stochastic gradient descent and variants thereof. Kronecker-factored Approximate Curvature (K-FAC), a natural gradient method, preconditions gradients with an efficient approximation of the Fisher Information Matrix to improve per-iteration progress when optimizing an objective function. Here we propose a scalable K-FAC algorithm and investigate K-FAC's applicability in large-scale deep neural network training. Specifically, we explore layer-wise distribution strategies, inverse-free second-order gradient evaluation, and dynamic K-FAC update decoupling, with the goal of preserving convergence while minimizing training time. We evaluate the convergence and scaling properties of our K-FAC gradient preconditioner, for image classification, object detection, and language modeling applications. In all applications, our implementation converges to baseline performance targets in 9-25% less time than the standard first-order optimizers on GPU clusters across a variety of scales.
Keyword:
Training
Parallel processing
Program processors
Convergence
Computational modeling
Data models
Deep learning
Optimization methods
neural networks
scalability
high-performance computing
期刊
IF:
6
论文数:
5.2K
被引数:
1.1W
机构
引用论文
Heterogeneous photooxidation of sulfur dioxide in the presence of airborne mineral dust particles
RSC Advances
IF0
Do Perceptions of Competence Mediate The Relationship Between Fundamental Motor Skill Proficiency and Physical Activity Levels of Children in Kindergarten?能力的感知是否可以介导幼儿园儿童的基本运动技能熟练程度与身体活动水平之间的关系?

