返回
A bio-inspired positional embedding network for transformer-based models
DOI:10.1016/j.neunet.2023.07.015.png)
摘要
En 中文
Owing to the progress of transformer-based networks, there have been significant improvements in the performance of vision models in recent years. However, there is further potential for improvement in positional embeddings that play a crucial role in distinguishing information across different positions. Based on the biological mechanisms of human visual pathways, we propose a positional embedding network that adaptively captures position information by modeling the dorsal pathway, which is responsible for spatial perception in human vision. Our proposed double-stream architecture leverages large zero-padding convolutions to learn local positional features and utilizes transformers to learn global features, effectively capturing the interaction between dorsal and ventral pathways. To evaluate the effectiveness of our method, we implemented experiments on various datasets, employing differentiated designs. Our statistical analysis demonstrates that the simple implementation significantly enhances image classification performance, and the observed trends demonstrate its biological plausibility.& COPY; 2023 Elsevier Ltd. All rights reserved.
Keyword:
Transformers
Dorsal pathway modeling
Image classification
Position embedding
Zero padding
期刊
IF:
6.3
论文数:
8.2K
被引数:
3.0W
机构
引用论文
Inflexibility of mental planning: A characteristic disorder with prefrontal lobe lesions?心理计划的僵化: 前额叶病变的特征性障碍?
Metal-organic chemical vapor deposition of indium selenide films using a single-source precursor使用单源前体对硒化铟薄膜进行金属有机化学气相沉积
V4 shape features for contour representation and object detection用于轮廓表示和目标检测的V4形状特征
NEURAL NETWORKS
IF6.3
Recycling of NdFeB Magnets from Electric Drive Motors of (Hybrid) Electric Vehicles从 (混合动力) 电动汽车的驱动电机中回收NdFeB磁体

