Return
An efficient Waveflow Vision Transformer for multi-scale NDE image processing
DOI:10.1080/10589759.2026.2658730.png)
Abstract
En 中文
In the industrial internet of things (IIoT) and non-destructive evaluation 4.0 (NDE 4.0), large-scale imagery data require capturing both global structural features and local anomalies. This presents challenges for feature extraction models in retaining local details and modelling global dependencies. The Vision Transformer (ViT) with Multi-Head Self-Attention (MHSA) is effective at modelling long-range dependencies but struggles with local feature capture and computational cost. To address these issues, we propose Waveflow, an integrated multi-head self-attention mechanism with a bottleneck structure. Different from existing wavelet-enhanced ViT variants that mainly focus on isolated wavelet decomposition in the attention layer, Waveflow introduces a co-design of wavelet-based frequency processing and bottleneck optimisation, achieving more efficient and comprehensive feature learning. In Waveflow, the proposed Wavelet Multi-Head Self-Attention (Wave-MHSA) is combined with a novel Wavelet Block Sparsity to subtly capture the complex details of global and local information in the spatial, frequency, and channel domains. Extensive NDT experiments show that Waveflow outperforms baseline methods, achieving mAP@0.5 of 81.3% on the NEU-DET dataset, 97.2% on the ASSD dataset, and 78.8% on the WCSD dataset.
Keywords:
Non-destructive non-destructive testing
vision transformers
multi-head self-attention
Discrete Wavelet Transform
Wavelet Block Sparsity
Journal
N
IF:
4.2
Papers:
1.8K
Citations:
2.1K
Organization
Cited Papers
WaveMamba-YOLO: Combining frequency awareness and state-space modeling for defect localization
PLoS One
IF2.6

