arrow
Return

Parallel Pipelined Architecture and Algorithm for Matrix Transposition Using Registers

delete2022-03-01
delete1
PRE
AI
B
Bo Zhang
Z
Zhenguo Ma
骆威 (Wei Luo) *
DOI:10.1109/TCSII.2021.3134710delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In this brief, we present a new algorithm and architecture for continuous-flow matrix transposition using registers. The algorithm supports P-parallel matrix transposition. The hardware architecture reaches the theoretical minimums in terms of latency and memory. It is composed of a group of identical cascaded basic swap circuits, whose stages are determined by the corresponding algorithm, and can be controlled via a set of counters. Compared with the state-of-the-art architecture, the proposed architecture supports matrices whose rows and columns are integer multiples of P. Here P can be arbitrary, including but not limited to power-of-two integers. Moreover, our results provide additional insight into continuous-flow non-square matrix transposition.
Keywords:
Computer architecture
Arrays
Hardware
Signal processing algorithms
Registers
Parallel processing
Circuits and systems
Pipelined algorithm
hardware architecture
continuous-flow
matrix transposition
parallel computing

Journal

I
IEEE Transactions on Circuits and Systems and Express Briefs
IF:
4.9
Papers:
8.8K
Citations:
2.5W

Organization

Z
zhejiang university
Scholars:
17.5W
Papers: 12.0W
Citations: 152