Return
Evaluating complex-valued transformers based on various architectural parameters for image classification
DOI:10.1016/j.ins.2026.123562.png)
Abstract
En 中文
Our research explores how three specific design settings shape the accuracy of complex-valued multi-head self-attention (MHSA) architectures on image classification tasks. The proposed pipeline employs complex-valued representations by converting real-valued images into the com plex domain; these are subsequently processed via patch embedding, augmented with a class token, and positional encoding, and finally passed through a complex-valued MHSA module. Tests on the standard MNIST and CIFAR-10 datasets revealed that each setting-the number of atten tion heads, the size of the query/key/value projections, and the image patch size-has a distinct and measurable impact on performance. These findings provide practical insights into how tuning hyperparameters in complex-valued transformers can influence image classification outcomes.
Keywords:
Complex-valued transformers
Complex-valued multi-head self-attention
Image classification
Journal
IF:
6.8
Papers:
540
Citations:
6.2W

