arrow
返回

Transformers in Vision: A Survey

delete2022-09-13
delete1.1K
delete
OA
AI
S
Salman Khan *
M
Muzammal Naseer
M
Munawar Hayat
S
Syed Waqas Zamir
F
Fahad Shahbaz Khan
M
Mubarak Shah
DOI:10.1145/3505244delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Astounding results from Transformer models on natural language tasks have intrigued the vision community to study their application to computer vision problems. Among their salient benefits, Transformers enable modeling long dependencies between input sequence elements and support parallel processing of sequence as compared to recurrent networks, e.g., Long short-term memory. Different from convolutional networks, Transformers require minimal inductive biases for their design and are naturally suited as set-functions. Furthermore, the straightforward design of Transformers allows processing multiple modalities (e.g., images, videos, text, and speech) using similar processing blocks and demonstrates excellent scalability to very large capacity networks and huge datasets. These strengths have led to exciting progress on a number of vision tasks using Transformer networks. This survey aims to provide a comprehensive overview of the Transformer models in the computer vision discipline. We start with an introduction to fundamental concepts behind the success of Transformers, i.e., self-attention, large-scale pre-training, and bidirectional feature encoding. We then cover extensive applications of transformers in vision including popular recognition tasks (e.g., image classification, object detection, action recognition, and segmentation), generative modeling, multi-modal tasks (e.g., visual-question answering, visual reasoning, and visual grounding), video processing (e.g., activity recognition, video forecasting), low-level vision (e.g., image super-resolution, image enhancement, and colorization), and three-dimensional analysis (e.g., point cloud classification and segmentation). We compare the respective advantages and limitations of popular techniques both in terms of architectural design and their experimental value. Finally, we provide an analysis on open research directions and possible future works. We hope this effort will ignite further interest in the community to solve current challenges toward the application of transformer models in computer vision.
Keyword:
Self-attention
transformers
bidirectional encoders
deep neural networks
convolutional networks
self-supervision
literature survey

期刊

ACM Computing Surveys 封面图
ACM Computing Surveys
IF:
28
论文数:
2.5K
被引数:
3.5W

机构

M
Monash University
学者数:
5.4W
论文数: 5.4W
被引数: 79
L
Linkoping University
学者数:
1.6W
论文数: 1.5W
被引数: 184
U
University of Central Florida
学者数:
8.7K
论文数: 6.8K
被引数: 1.4W
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
In Situ Hybridization Studies for Virai Nucleic Acids in Heart and Lung Allograft Biopsies
err1990-05-01
err0
PREAI
errLawrence M. Weiss; Lucile A. Movahed; Gerald J. Berry; Margaret E. Billingham
err分享
err收藏
Using SEPIC Topology for Improving Power Factor in Distributed Power Supply Systems
err2015-09-22
err0
PREAI
errJ. Sebastián; J. Uceda; J.A. Cobos; J. Arau
err分享
err收藏
Oxidative damage to fibronectin
err1991-02-01
err0
PREAI
errMargret C.M. Vissers; Christine C. Winterbourn
err分享
err收藏
Obtention and characterization of primary astrocyte and microglial cultures from adult monkey brains
err1997-09-01
err0
errOAAI
errG. Guillemin; F.D. Boussin; J. Croitoru; M. Franck-Duchenne; R. Le Grand; F. Lazarini; D. Dormont
err分享
err收藏
err分享
err收藏
学者 查看更多内容