1
Return

Multi-DCU parallelization of the FW–H solver for aeroacoustic simulations

delete2026-06-24
delete0
PRE
AI
Q
Qixing Wang
M
Mingyu Shao
H
Hanbo Jiang *
DOI:10.1016/j.cpc.2026.110284delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The Ffowcs Williams–Hawkings (FW–H) acoustic analogy is widely used for noise prediction, but large integration surfaces and long source-time histories can make the post-processing computationally expensive. To address this issue, we present a Message Passing Interface (MPI)-based multi-node implementation of the FW–H solver for platforms equipped with Deep Computing Units (DCUs), i.e., GPU-class accelerators. In this framework, host CPUs handle data preparation, inter-node communication, and I/O, while the computationally dominant FW–H integration is offloaded to the DCUs. The nested loops over observers, time/frequency samples, and surface elements are mapped onto a three-dimensional thread grid, and distributed domain decomposition is adopted to reduce the memory footprint on each rank. The implementation is validated on canonical two-dimensional and three-dimensional benchmarks, with mean relative errors in far-field directivity below 0.15% and 0.70%, respectively. For a representative 2D frequency-domain benchmark, a single Hygon DCU achieves an end-to-end speedup of up to 21.2 ×  over an optimized 32-core Hygon C86 7185 CPU implementation, and up to 708 ×  for the FW–H integration kernel. In this 2D case, the CPU implementation scales nearly linearly at low to moderate core counts but shows signs of saturation near 32 cores, whereas for the large 3D time-domain benchmark, both CPU and DCU implementations exhibit robust multi-node scalability up to 128 CPU cores and 128 DCUs, with the DCU path showing near-linear scaling and, in some configurations, superlinear speedup. Scalability analysis indicates that the CPU saturation trend is mainly associated with increasing MPI overhead and intra-node resource contention, whereas the superlinear DCU speedup arises primarily from reduced chunking and lower host–device transfer overhead as the per-device workload decreases. A rated-power-based estimate also suggests lower energy-to-solution for the DCU configurations than for the corresponding CPU configurations. Overall, the proposed multi-DCU implementation provides an accurate, efficient, and scalable framework for FW–H-based aeroacoustic prediction.

Journal

Computer Physics Communications cover
Computer Physics Communications
IF:
3.4
Papers:
1.2W
Citations:
3.7W

Organization

S
shanghai jiao tong university
Scholars:
15.1W
Papers: 11.5W
Citations: 159
P
peking university
Scholars:
11.5W
Papers: 8.6W
Citations: 146
E
Eastern Institute of Technology
Scholars:
551
Papers: 395
Citations: 825
Cited Papers

Cited Papers

Citing Papers

Citing Papers