arrow
返回

An optimized second order stochastic learning algorithm for neural network training

delete2016-04-01
delete28
PRE
AI
S
Shan Sung Liew *
M
Mohamed Khalil-Hani
R
Rabia Bakhteri
DOI:10.1016/j.neucom.2015.12.076delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This paper proposes an improved stochastic second order learning algorithm for supervised neural network training. The proposed algorithm, named bounded stochastic diagonal Levenberg-Marquardt (B-SDLM), utilizes both gradient and curvature information to achieve fast convergence while requiring only minimal computational overhead than the stochastic gradieht descent (SGD)method. B-SDLM has only a single hyperparameter as opposed to most other learning algorithms that suffer from the hyperparameter overfitting problem due to having more hyperparameters to be tuned. Experiments using the multilayer perceptron (MLP) and convolutional neural network (CNN) models have shown that B-SDLM outperforms other learning algorithms with regard to the classification accuracies and computational efficiency (about 5.3% faster than SGD on the mnist-rot-hg-img database). It can classify all testing samples correctly on the face recognition case study based on AR Purdue database. In addition, experiments on handwritten digit classification case studies show that significant improvements of 19.6% on MNIST database and 17.5% on mnist-rot-bg-img database can be achieved in terms of the testing misclassification error rates (MCRs). The computationally expensive Hessian calculations are kept to a minimum by using just 0.05% of the training samples in its estimation or updating the learning rates once per two training epochs, while maintaining or even achieving lower testing MCRs. It is also shown that B-SDLM works well in the mini-batch learning mode, and we are able to achieve 3.32x performance speedup when deploying the proposed algorithm in a distributed learning environment with a quad-core processor. (C) 2016 Elsevier B.V. All rights reserved.
Keyword:
Stochastic diagonal Levenberg-Marquardt
Fast convergence
Hyperparameter overfitting
Computational efficiency
Distributed machine learning
Convolutional neural network

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

U
Universiti Teknologi Malaysia
学者数:
1.4W
论文数: 1.1W
被引数: 85
引用论文

引用论文

Learning for 3D understanding
err2015-03-01
err2
PREAI
errGao, Yue; Ji, Rongrong; Liu, Wei; Dai, Qionghai
err分享
err收藏
Multimodal Deep Autoencoder for Human Pose Recovery
err2015-12-01
err520
PREAI
errHong, Chaoqun; Yu, Jun; Wan, Jian; Tao, Dacheng; Wang, Meng
err分享
err收藏
Dictionary-Based Face Recognition Under Variable Lighting and Pose
err2012-06-01
err86
PREAI
errPatel, Vishal M.; Wu, Tao; Biswas, Soma; Phillips, P. Jonathon; Chellappa, Rama
err分享
err收藏
err分享
err收藏
err分享
err收藏
Heterotic geometry without isometries
err2005-10-20
err0
errOAAI
errAlexei P Isaev; Osvaldo P Santillan
err分享
err收藏
学者 查看更多内容