Return
Parallel and Pipelined BRAM-Based Matrix Transposition for 6G
J
C
X
M
DOI:10.1109/TCSII.2025.3584052.png)
Abstract
En 中文
In this brief, we present a parallel and pipelined algorithm for BRAM-based matrix transposition, along with its corresponding architecture, optimized specifically to meet the stringent throughput and latency demands of 6G. The architecture utilizes a novel address mapping algorithm, which exploits the coprimality between memory parameters to achieve conflict-free parallel access via a simple yet efficient prime-modulo addressing scheme.The architecture achieves conflict-free parallel memory access on BRAM, significantly improving parallelism and enhancing throughput. More importantly, by adopting a ping-pong buffering scheme, it enables fully pipelined and highly parallel matrix transposition, primarily targeting low-latency and high-throughput tasks in 6G. Experimental results show that, compared with existing implementations supporting similar matrix sizes, the architecture in this brief increases throughput significantly from 0.8 GB/s to 25.6 GB/s under a latency of 0.08ms.
Keywords:
Matrix transposition
BRAM
6G
pipelined architecture
parallel memory access
Journal
I
IF:
4.9
Papers:
8.8K
Citations:
2.5W
