arrow
Return

Programming parallel dense matrix factorizations and inversion for new-generation NUMA architectures

delete2023-05-01
delete1
delete
OA
AI
S
Sandra Catalán
F
Francisco D. Igual
J
José R. Herrero
R
Rafael Rodríguez‐Sánchez
E
Enrique S. Quintana–Ort́ı *
DOI:10.1016/j.jpdc.2023.01.004delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
We propose a methodology to address the programmability issues derived from the emergence of newgeneration shared-memory NUMA architectures. For this purpose, we employ dense matrix factorizations and matrix inversion (DMFI) as a use case, and we target two modern architectures (AMD Rome and Huawei Kunpeng 920) that exhibit configurable NUMA topologies. Our methodology pursues performance portability across different NUMA configurations by proposing multi-domain implementations for DMFI plus a hybrid task- and loop-level parallelization that configures multi-threaded executions to fix core-todata binding, exploiting locality at the expense of minor code modifications. In addition, we introduce a generalization of the multi-domain implementations for DMFI that offers support for virtually any NUMA topology in present and future architectures. Our experimentation on the two target architectures for three representative dense linear algebra operations validates the proposal, reveals insights on the necessity of adapting both the codes and their execution to improve data access locality, and reports performance across architectures and inter- and intra-socket NUMA configurations competitive with state-of-the-art message-passing implementations, maintaining the ease of development usually associated with shared-memory programming. (c) 2023 The Author(s). Published by Elsevier Inc. This is an open access article under the CC BY-NC-ND
Keywords:
NUMA architectures
Chiplets
Dense linear algebra
Shared memory programming
Portability
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

C
Complutense University of Madrid
Scholars:
2.6W
Papers: 2.2W
Citations: 31
U
Universitat Politecnica de Valencia
Scholars:
1.5W
Papers: 1.4W
Citations: 18
U
universitat politecnica de catalunya
Scholars:
1.9W
Papers: 1.6W
Citations: 17
researcher View more organizations