arrow
返回

Large Linear Classification When Data Cannot Fit in Memory

delete2012-02-01
delete57
delete
OA
AI
H
Hsiang‐Fu Yu *
C
Cho‐Jui Hsieh
K
Kai‐Wei Chang
C
Chih‐Jen Lin
DOI:10.1145/2086737.2086743delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Recent advances in linear classification have shown that for applications such as document classification, the training process can be extremely efficient. However, most of the existing training methods are designed by assuming that data can be stored in the computer memory. These methods cannot be easily applied to data larger than the memory capacity due to the random access to the disk. We propose and analyze a block minimization framework for data larger than the memory size. At each step a block of data is loaded from the disk and handled by certain learning methods. We investigate two implementations of the proposed framework for primal and dual SVMs, respectively. Because data cannot fit in memory, many design considerations are very different from those for traditional algorithms. We discuss and compare with existing approaches that are able to handle data larger than memory. Experiments using data sets 20 times larger than the memory demonstrate the effectiveness of the proposed method.
Keyword:
Block minimization methods
large-scale learning
linear classification
support vector machines
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

ACM Transactions on Knowledge Discovery from Data 封面图
ACM Transactions on Knowledge Discovery from Data
IF:
4.8
论文数:
1.3K
被引数:
4.4K

机构

N
National Taiwan University
学者数:
4.7W
论文数: 4.2W
被引数: 3.6W
引用论文

引用论文

Percentage White: A New Feature for Ultrasound Classification of Plaque Echogenicity in Carotid Artery Atherosclerosis
err2010-02-01
err0
PREAI
errUlrica Prahl; Peter Holdfeldt; Göran Bergström; Björn Fagerberg; Johannes Hulthe; Tomas Gustavsson
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容