arrow
Return

Large-scale distributed linear algebra with tensor processing units

delete2022-08-08
delete11
delete
OA
AI
A
Adam G. M. Lewis *
J
Jackson Beall
M
Martin Ganahl
M
Markus Hauru
S
Shrestha Basu Mallick
V
Vidal, Guifre
DOI:10.1073/pnas.2122762119delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
We have repurposed Google tensor processing units (TPUs), application-specific chips developed for machine learning, into large-scale dense linear algebra supercomputers. The TPUs' fast intercore interconnects (ICIs), physically two-dimensional network topology, and high-bandwidth memory (HBM) permit distributed matrix multiplication algorithms to rapidly become computationally bound. In this regime, the matrix-multiply units (MXUs) dominate the runtime, yielding impressive scaling, performance, and raw size: Operating in float32 precision, a full 2,048-core pod of third-generation TPUs can multiply two matrices with linear size N = 2(20) = 1,048,576 in about 2 min. Via curated algorithms emphasizing large, single-core matrix multiplications, other tasks in dense linear algebra can similarly scale. As examples, we present 1) QR decomposition; 2) resolution of linear systems; and 3) the computation of matrix functions by polynomial iteration, demonstrated by the matrix polar factorization.
Keywords:
TPUs
scientific computation
linear algebra
distributed computing
ASICs
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

P
Proceedings of the National Academy of Sciences of the United States of America
IF:
9.1
Papers:
10.8W
Citations:
73.5W

Organization

G
Google Incorporated
Scholars:
3.5K
Papers: 1.8K
Citations: 8