arrow
Return

Large-Scale Parallel Embedded Computing With Improved-MPI in Off-Chip Distributed Clusters

delete2024-11-01
delete0
PRE
AI
X
Xile Wei
H
Hengyi Wei
卢梅丽 (Meili Lu)
张镇 cover
张镇 (Zhen Zhang)
S
Siyuan Chang *
S
Shunqi Zeng
F
Fei Wang
M
Minghui Chen
DOI:10.1109/TII.2024.3423330delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Distributed architecture is expected to be an effective solution for large-scale edge computing tasks in terminal devices. However, it remains a great challenge to resolve the conflict between parallel efficiency and constrained physical resources in a specific embedded structure. This article proposes a universal scalable off-chip parallel computing architecture to maximize the computing efficiency for distributed embedded computing clusters. This architecture is based on an improved Message Passing Interface (Improved-MPI). To address the limited communication speed in embedded environments, a multilevel communication mechanism is employed to alleviate the communication pressure on nodes. By flexibly allocating computing tasks, efficient utilization of every embedded cluster node is ensured, while also solving the problem of single point of failure. In addition, to overcome the challenge of limited RAM in embedded devices, the architecture utilizes the interleaved memory initialization mechanism to run larger computing tasks. Based on this architecture, a specific embedded cluster platform is constructed using the RK3399 board. Various large-scale tasks are deployed on this platform to validate the performance of the architecture. First, a large-scale randomly connected neural network is executed, which serves to verify the architecture's outstanding computational performance and communication capability. Secondly, a functional model of Small-World Spiking Neural Network is constructed, achieving real-time and efficient digital speech recognition. Finally, the implementation of Large Language Models demonstrates that the embedded clusters can achieve performance comparable to modern computers.
Keywords:
Distributed computing
embedded system
large language model (LLM)
message passing interface (MPI)
neural networks
Distributed computing
embedded system
large language model (LLM)
message passing interface (MPI)
neural networks

Journal

IEEE Transactions on Industrial Informatics cover
IEEE Transactions on Industrial Informatics
IF:
9.9
Papers:
8.3K
Citations:
6.0W

Organization

T
tianjin university
Scholars:
7.9W
Papers: 5.7W
Citations: 88
T
tianjin university of technology & education
Scholars:
769
Papers: 638
Citations: 0