arrow
返回

Multi-granularity sequence generation for hierarchical image classification

delete2024-04-01
delete0
delete
OA
AI
X
Xinda Liu
L
Lili Wang *
DOI:10.1007/s41095-022-0332-2delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Hierarchical multi-granularity image classification is a challenging task that aims to tag each given image with multiple granularity labels simultaneously. Existing methods tend to overlook that different image regions contribute differently to label prediction at different granularities, and also insufficiently consider relationships between the hierarchical multi-granularity labels. We introduce a sequence-to-sequence mechanism to overcome these two problems and propose a multi-granularity sequence generation (MGSG) approach for the hierarchical multi-granularity image classification task. Specifically, we introduce a transformer architecture to encode the image into visual representation sequences. Next, we traverse the taxonomic tree and organize the multi-granularity labels into sequences, and vectorize them and add positional information. The proposed multi-granularity sequence generation method builds a decoder that takes visual representation sequences and semantic label embedding as inputs, and outputs the predicted multi-granularity label sequence. The decoder models dependencies and correlations between multi-granularity labels through a masked multi-head self-attention mechanism, and relates visual information to the semantic label information through a cross-modality attention mechanism. In this way, the proposed method preserves the relationships between labels at different granularity levels and takes into account the influence of different image regions on labels with different granularities. Evaluations on six public benchmarks qualitatively and quantitatively demonstrate the advantages of the proposed method. Our project is available at https://github.com/liuxindazz/mgsg.
Keyword:
hierarchical multi-granularity classification
vision and text transformer
sequence generation
fine-grained image recognition
cross-modality attention

期刊

Computational Visual Media 封面图
Computational Visual Media
IF:
18.3
论文数:
315
被引数:
2.6K

机构

B
Beihang University
学者数:
5.2W
论文数: 4.1W
被引数: 37
引用论文

引用论文

err分享
err收藏
High prevalence of malaria in a non-endemic setting among febrile episodes in travellers and migrants coming from endemic areas: a retrospective analysis of a 2013–2018 cohort
err2021-11-27
err0
errOAAI
errAlejandro Garcia-Ruiz de Morales; Covadonga Morcate; Elena Isaba-Ares; Ramon Perez-Tanoira; Jose A. Perez-Molina
err分享
err收藏
Transformers in computational visual media: A survey
err2022-03-01
err85
errOAAI
errXu, Yifan; Wei, Huapeng; Lin, Minxuan; Deng, Yingying; Sheng, Kekai; Zhang, Mengdan; Tang, Fan; Dong, Weiming; Huang, Feiyue; Xu, Changsheng
err分享
err收藏
Bacterial and algal markers in sedimentary organic matter deposited under natural sulphurization conditions (Lorca Basin, Murcia, Spain)
err1997-05-01
err0
PREAI
errMarie Russell; Joan O. Grimalt; Walter A. Hartgers; Conxita Taberner; Jean Marie Rouchy
err分享
err收藏
学者 查看更多内容