arrow
Return

Fast block distributed CUDA implementation of the Hungarian algorithm

delete2019-08-01
delete17
PRE
AI
P
Paulo A. C. Lopes *
S
Satyendra Singh Yadav
A
Aleksandar Ilić
S
Sarat Kumar Patra
DOI:10.1016/j.jpdc.2019.03.014delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The Hungarian algorithm solves the linear assignment problem in polynomial time. A GPU/CUDA implementation of this algorithm is proposed. GPUs are massive parallel machines. In this implementation, the alternating path search phase of the algorithm is distributed by several blocks in a way to minimize global device synchronization. This phase is very important and has a big contribution to the execution time. Other advanced features also implemented are: parallel graph traversal; the parallel detection of multiple alternating paths in a single iteration; a simplified and fast matrix compression that stores the zeros of the slack matrix, resulting in very fast graph traversal; highly optimized reductions for the initial slack matrix calculation and update. This results in a fast implementation for moderate size problems. (C) 2019 Elsevier Inc. All rights reserved.
Keywords:
Hungarian algorithm
Linear assignment problem
GPU
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

U
universidade de lisboa
Scholars:
3.4W
Papers: 3.1W
Citations: 29
N
national institute of technology (nit system)
Scholars:
4.0W
Papers: 3.7W
Citations: 31
I
inesc-id
Scholars:
636
Papers: 504
Citations: 0
researcher View more organizations