arrow
Return

Fast and Communication-Efficient Algorithm for Distributed Support Vector Machine Training

delete2019-05-01
delete17
PRE
AI
J
Jyotikrishna Dass *
R
Rabi Mahapatra
DOI:10.1109/TPDS.2018.2879950delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Support Vector Machines (SVM) are widely used as supervised learning models to solve the classification problem in machine learning. Training SVMs for large datasets is an extremely challenging task due to excessive storage and computational requirements. To tackle so-called big data problems, one needs to design scalable distributed algorithms to parallelize the model training and to develop efficient implementations of these algorithms. In this paper, we propose a distributed algorithm for SVM training that is scalable and communication-efficient. The algorithm uses a compact representation of the kernel matrix, which is based on the QR decomposition of low-rank approximations, to reduce both computation and storage requirements for the training stage. This is accompanied by considerable reduction in communication required for a distributed implementation of the algorithm. Experiments on benchmark data sets with up to five million samples demonstrate negligible communication overhead and scalability on up to 64 cores. Execution times are vast improvements over other widely used packages. Furthermore, the proposed algorithm has linear time complexity with respect to the number of samples making it ideal for SVM training on decentralized environments such as smart embedded systems and edge-based internet of things, IoT.
Keywords:
Machine learning
support vector machines
classification algorithms
parallel programming
distributed computing
message passing
quadratic programming
iterative algorithms
optimization
multicore processing
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

T
Texas A&M University System
Scholars:
4.4W
Papers: 4.0W
Citations: 4.0K
Cited Papers

Cited Papers

Percentage White: A New Feature for Ultrasound Classification of Plaque Echogenicity in Carotid Artery Atherosclerosis
err2010-02-01
err0
PREAI
errUlrica Prahl; Peter Holdfeldt; Göran Bergström; Björn Fagerberg; Johannes Hulthe; Tomas Gustavsson
errShare
errSave
err
IF0
err
err0
errOAAI
err
errShare
errSave
COMMUNICATION-OPTIMAL PARALLEL AND SEQUENTIAL QR AND LU FACTORIZATIONS
err2012-01-01
err237
errOAAI
errDemmel, James; Grigori, Laura; Hoemmen, Mark; Langou, Julien
errShare
errSave
Stereotyped: Investigating Gender in Introductory Science Courses
err2013-03-01
err0
errOAAI
errShanda Lauer; Jennifer Momsen; Erika Offerdahl; Mila Kryjevskaia; Warren Christensen; Lisa Montplaisir
errShare
errSave
Base metal Co-fired (Na,K)NbO3 structures with enhanced piezoelectric performance
err2014-02-21
err0
PREAI
errCheng Liu; Peng Liu; Keisuke Kobayashi; Clive A. Randall
errShare
errSave
no more