arrow
Return

Knowledge Amalgamation for Object Detection With Transformers

delete2023-01-01
delete8
delete
OA
AI
H
Haofei Zhang
F
Feng Mao
M
Mengqi Xue
G
Gongfan Fang
冯尊磊 (Zunlei Feng)
宋杰 (Jie Song)
宋明黎 (Mingli Song) *
DOI:10.1109/TIP.2023.3263105delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Knowledge amalgamation (KA) is a novel deep model reusing task aiming to transfer knowledge from several well-trained teachers to a multi-talented and compact student. Currently, most of these approaches are tailored for convolutional neural networks (CNNs). However, there is a tendency that Transformers, with a completely different architecture, are starting to challenge the domination of CNNs in many computer vision tasks. Nevertheless, directly applying the previous KA methods to Transformers leads to severe performance degradation. In this work, we explore a more effective KA scheme for Transformer-based object detection models. Specifically, considering the architecture characteristics of Transformers, we propose to dissolve the KA into two aspects: sequence-level amalgamation (SA) and task-level amalgamation (TA). In particular, a hint is generated within the sequence-level amalgamation by concatenating teacher sequences instead of redundantly aggregating them to a fixed-size one as previous KA approaches. Besides, the student learns heterogeneous detection tasks through soft targets with efficiency in the task-level amalgamation. Extensive experiments on PASCAL VOC and COCO have unfolded that the sequence-level amalgamation significantly boosts the performance of students, while the previous methods impair the students. Moreover, the Transformer-based students excel in learning amalgamated knowledge, as they have mastered heterogeneous detection tasks rapidly and achieved superior or at least comparable performance to those of the teachers in their specializations.
Keywords:
Transformers
Task analysis
Object detection
Detectors
Training
Computer architecture
Feature extraction
Model reusing
knowledge amalgamation
knowledge distillation
object detection
vision transformers

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

A
alibaba group
Scholars:
1.1K
Papers: 789
Citations: 0
H
Hangzhou City University
Scholars:
2.2K
Papers: 2.0K
Citations: 1.0K
N
National University of Singapore
Scholars:
7.5W
Papers: 6.4W
Citations: 11.4W
Z
zhejiang university
Scholars:
17.4W
Papers: 12.0W
Citations: 152
researcher View more organizations