arrow
Return

Efficient and Distributed Generalized Canonical Correlation Analysis for Big Multiview Data

delete2019-12-01
delete16
PRE
AI
X
Xiao Fu *
K
Kejun Huang
E
Evangelos E. Papalexakis
H
Hyun Ah Song
T
Talukdar, Partha
N
Nicholas D. Sidiropoulos
C
Christos Faloutsos
T
Tom M. Mitchell
DOI:10.1109/TKDE.2018.2875908delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Generalized canonical correlation analysis (GCCA) integrates information from data samples that are acquired at multiple feature spaces (or 'views') to produce low-dimensional representations-which is an extension of classical two-view CCA. Since the 1960s, (G)CCA has attracted much attention in statistics, machine learning, and data mining because of its importance in data analytics. Despite these efforts, the existing GCCA algorithms have serious complexity issues. The memory and computational complexities of the existing algorithms usually grow as a quadratic and cubic function of the problem dimension (the number of samples / features), respectively-e.g., handling views with approximate to 1,000 features using such algorithms already occupies approximate to 10(6) memory and the per-iteration complexity is approximate to 10(9) flops-which makes it hard to push these methods much further. To circumvent such difficulties, we first propose a GCCA algorithm whose memory and computational costs scale linearly in the problem dimension and the number of nonzero data elements, respectively. Consequently, the proposed algorithm can easily handle very large sparse views whose sample and feature dimensions both exceed approximate to 100,000. Our second contribution lies in proposing two distributed algorithms for GCCA, which compute the canonical components of different views in parallel and thus can further reduce the runtime significantly if multiple computing agents are available. We provide detailed convergence analyses of the proposed algorithms and show that all the large-scale GCCA algorithms converge to a Karush-Kuhn-Tucker (KKT) point at least sublinearly. Judiciously designed synthetic and real-data experiments are employed to showcase the effectiveness of the proposed algorithms.
Keywords:
Distributed algorithms
Sparse matrices
Correlation
Machine learning algorithms
Electronic mail
Data mining
Machine learning
Generalized canonical correlation analysis
multiview learning
multilingual word embedding
distributed GCCA
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Knowledge and Data Engineering cover
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
Papers:
6.7K
Citations:
3.2W

Organization

U
University of Florida
Scholars:
4.0W
Papers: 3.1W
Citations: 6.6W
State University System of Florida cover
State University System of Florida
Scholars:
12.6W
Papers: 10.8W
Citations: 130
U
university of california riverside
Scholars:
1.0W
Papers: 8.2K
Citations: 16
O
Oregon State University
Scholars:
1.7W
Papers: 1.5W
Citations: 2.4W
I
indian institute of science (iisc) - bangalore
Scholars:
1.4W
Papers: 1.4W
Citations: 11
University of California System cover
University of California System
Scholars:
37.2W
Papers: 33.6W
Citations: 6.6K
researcher View more organizations