返回
Fine-grained semantic oriented embedding set alignment for text-based person search
DOI:10.1016/j.imavis.2024.105309.png)
摘要
En 中文
Text-based person search aims to retrieve images of a person that are highly semantically relevant to a given textual description. The difficulty of this retrieval task is modality heterogeneity and fine-grained matching. Most existing methods only consider the alignment using global features, ignoring the fine-grained matching problem. The cross-modal attention interactions are popularly used for image patches and text markers for direct alignment. However, cross-modal attention may cause a huge overhead in the reasoning stage and cannot be applied in actual scenarios. In addition, it is unreasonable to perform patch-token alignment, since image patches and text tokens do not have complete semantic information. This paper proposes an Embedding Set Alignment (ESA) module for fine-grained alignment. The module can preserve fine-grained semantic information by merging token-level features into embedding sets. The ESA module benefits from pre-trained cross-modal large models, and it can be combined with the backbone non-intrusively and trained in an end-to-end manner. In addition, an Adaptive Semantic Margin (ASM) loss is designed to describe the alignment of embedding sets, instead of adapting a loss function with a fixed margin. Extensive experiments demonstrate that our proposed fine-grained semantic embedding set alignment method achieves state-of-the-art performance on three popular benchmark datasets, surpassing the previous best methods.
Keyword:
Text-based person search
Cross modal
Person re-identification
Fine-grained
期刊
IF:
4.2
论文数:
4.1K
被引数:
6.7K
机构
暂无机构信息
引用论文
Feature alignment via mutual mapping for few-shot fine-grained visual classification通过相互映射进行特征对齐,用于少镜头细粒度视觉分类
VLCDoC: Vision-Language contrastive pre-training model for cross-Modal document classificationVLCDoC: 跨模态文档分类的视觉语言对比预训练模型
PATTERN RECOGNITION
IF7.6
F-SCP: An automatic prompt generation method for specific classes based on visual language pre-training models
PATTERN RECOGNITION
IF7.6

