返回
Dynamic attention guider network
DOI:10.1007/s00607-024-01328-4.png)
摘要
En 中文
Hybrid networks, benefiting from both CNNs and Transformers architectures, exhibit stronger feature extraction capabilities compared to standalone CNNs or Transformers. However, in hybrid networks, the lack of attention in CNNs or insufficient refinement in attention mechanisms hinder the highlighting of target regions. Additionally, the computational cost of self-attention in Transformers poses a challenge to further improving network performance. To address these issues, we propose a novel point-to-point Dynamic Attention Guider(DAG) that dynamically generates multi-scale large receptive field attention to guide CNN networks to focus on target regions. Building upon DAG, we introduce a new hybrid network called the Dynamic Attention Guider Network (DAGN), which effectively combines Dynamic Attention Guider Block (DAGB) modules with Transformers to alleviate the computational cost of self-attention in processing high-resolution input images. Extensive experiments demonstrate that the proposed network outperforms existing state-of-the-art models across various downstream tasks. Specifically, the network achieves a Top-1 classification accuracy of 88.3% on ImageNet1k. For object detection and instance segmentation on COCO, it respectively surpasses the best FocalNet-T model by 1.6 APb\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$AP<^>b$$\end{document} and 1.5 APm\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$AP<^>m$$\end{document}, while achieving a top performance of 48.2% in semantic segmentation on ADE20K.
Keyword:
Hybrid networks
Multi-scale
Multi-path
Attention
Backbone
期刊
C
IF:
2.8
论文数:
2.3K
被引数:
3.5K
机构
引用论文
Caw’s Walking State Recognition Based on Accelerometers and Gyroscopes Installed on Ear-Tags and Collar-Tags基于安装在耳标和项圈上的加速度计和陀螺仪的Caw步行状态识别
Inflexibility of mental planning: A characteristic disorder with prefrontal lobe lesions?心理计划的僵化: 前额叶病变的特征性障碍?
Recent Development of Dual-Dictionary Learning Approach in Medical Image Analysis and Reconstruction
GhostFormer: Efficiently amalgamated CNN-transformer architecture for object detection
PATTERN RECOGNITION
IF7.6

