返回
Gradient Descent Learning With Floats
DOI:10.1109/TCYB.2020.2997399.png)
摘要
En 中文
The gradient learning descent method is the main workhorse of training tasks in artificial intelligence and machine-learning research. Current theoretical studies of gradient descent only use the continuous domains, which is unreal since electronic computers use the float point numbers to store and deal with data. Although existing results are sufficient for the extremely tiny errors in high-precision machines, they need to be improved for low-precision cases. This article presents an understanding of the learning algorithm in computers with floats. The performances of three gradient descents with the floating domain are investigated when the objective function is smooth. When the function is assumed to have the PL condition, the convergence speed can be improved. We proved that for floating gradient descent to obtain an error with is an element of, the iteration is O(1/is an element of) for the general smooth case, and O(ln(1/is an element of)) for the PL case. But is an element of should be larger than the s-bit machine epsilon delta(s) in the deterministic case, that is, is an element of >= Omega (delta(s)), while is an element of >= Omega (root delta(s)) for the stochastic case. Floating stochastic and sign gradient descents can both output an is an element of noised result in O(1/is an element of(2)) iterations.
Keyword:
Complexity
convergence
floats
gradient descent
learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
10.5
论文数:
1.1W
被引数:
5.0W
机构
引用论文
ICT for informal workers in Sub-Saharan Africa: Systematic review and analysis撒哈拉以南非洲非正规工人的信通技术: 系统回顾和分析

