返回
TripleFormer: improving transformer-based image classification method using multiple self-attention inputs
DOI:10.1007/s00371-024-03294-6.png)
摘要
En 中文
Transformer network structures have significantly improved the performance of computer vision (CV) tasks. However, due to the restriction of taking high-dimensional tokens as model inputs and the singularity of tokenization method, it is computationally intensive and easily deteriorates local fine-grained features when recognizing images based on Transformer architecture. In this paper, we incorporate a novel triple self-attention mechanism as a single encoder block and integrate it with the Transformer structure to introduce a new model, namely TripleFormer. Firstly, modeling information inside the window rather than the entire image, we present two sequence inputs from orthogonal perspectives, which are named Patch Attention In Spatial (PAIS) and Patch Attention In Channel (PAIC) for capturing more detailed features. We further partition the multi-channel of one feature map along spatial dimensions and compute attention belonging to same channels. In this way, the proposed encoder block incorporates local feature extraction and long-range visual dependencies to boost the feature learning capability. Finally, experiments on ImageNet-1K and CIFAR100 datasets exhibit the superiority of our proposed models as compared to other methods in terms of lower FLOPs and complexity while maintaining similar accuracy. In addition, our models demonstrate competitive performance on small-scale datasets in comparison to other pure Transformer models.
Keyword:
Image classification
Convolution neural networks
Visual transformer
Deep learning
期刊
IF:
2.9
论文数:
4.6K
被引数:
6.5K
机构
引用论文
Developing Wind and/or Solar Powered Crop Irrigation Systems for the Great Plains为大平原开发风能和/或太阳能作物灌溉系统
Molecular and Pharmacological Characterization of GABAA Receptors in the Rat Pituitary大鼠垂体中GABAA 受体的分子和药理学特征
Inflexibility of mental planning: A characteristic disorder with prefrontal lobe lesions?心理计划的僵化: 前额叶病变的特征性障碍?
Comparative study of layer by layer assembled multilayer films based on graphene oxide and reduced graphene oxide on flexible polyurethane foam: flame retardant and smoke suppression properties
RSC Advances
IF0
Thermal comfort, perceived air quality, and cognitive performance when personally controlled air movement is used by tropically acclimatized persons当热带适应的人使用个人控制的空气运动时,热舒适性,感知的空气质量和认知表现
Indoor Air
IF0

