arrow
返回

Conditional Feature Learning Based Transformer for Text-Based Person Search

delete2022-01-01
delete9
PRE
AI
C
Chenyang Gao
G
Guanyu Cai
X
Xinyang Jiang
F
Feng Zheng *
张
张军 (Jun Zhang)
Y
Yifei Gong
F
Fangzhou Lin
X
Xing Sun
白
白翔 (Xiang Bai)
DOI:10.1109/TIP.2022.3205216delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Text-based person search aims at retrieving the target person in an image gallery using a descriptive sentence of that person. The core of this task is to calculate a similarity score between the pedestrian image and description, which requires inferring the complex latent correspondence between image sub-regions and textual phrases at different scales. Transformer is an intuitive way to model the complex alignment by its self-attention mechanism. Most previous Transformer-based methods simply concatenate image region features and text features as input and learn a cross-modal representation in a brute force manner. Such weakly supervised learning approaches fail to explicitly build alignment between image region features and text features, causing an inferior feature distribution. In this paper, we present CFLT, Conditional Feature Learning based Transformer. It maps the sub-regions and phrases into a unified latent space and explicitly aligns them by constructing conditional embeddings where the feature of data from one modality is dynamically adjusted based on the data from the other modality. The output of our CFLT is a set of similarity scores for each sub-region or phrase rather than a cross-modal representation. Furthermore, we propose a simple and effective multi-modal re-ranking method named Re-ranking scheme by Visual Conditional Feature (RVCF). Benefit from the visual conditional feature and better feature distribution in our CFLT, the proposed RVCF achieves significant performance improvement. Experimental results show that our CFLT outperforms the state-of-the-art methods by 7.03% in terms of top-1 accuracy and 5.01% in terms of top-5 accuracy on the text-based person search dataset.
Keyword:
Feature extraction
Transformers
Visualization
Task analysis
Representation learning
Electronic mail
Data mining
Text-based person search
transformer
multi-granularity image-text alignments

期刊

IEEE Transactions on Image Processing 封面图
IEEE Transactions on Image Processing
IF:
13.7
论文数:
1.0W
被引数:
8.4W

机构

T
tohoku university
学者数:
4.3W
论文数: 3.6W
被引数: 31
T
Tencent
学者数:
1.1K
论文数: 898
被引数: 5
引用论文

引用论文

Dual-path Convolutional Image-Text Embeddings with Instance Loss
err2020-05-22
err345
errOAAI
errZheng, Zhedong; Zheng, Liang; Garrett, Michael; Yang, Yi; Xu, Mingliang; Shen, Yi-Dong
err分享
err收藏
Study of molecular interactions with 13C DNP-NMR
err2010-03-01
err0
PREAI
errMathilde H. Lerche; Sebastian Meier; Pernille R. Jensen; Herbert Baumann; Bent O. Petersen; Magnus Karlsson; Jens Ø. Duus; Jan H. Ardenkjær-Larsen
err分享
err收藏
Metagenomic assessment of adventitious viruses in commercial bovine sera
err2017-05-01
err0
PREAI
errKathy Toohey-Kurth; Samuel D. Sibley; Tony L. Goldberg
err分享
err收藏
err分享
err收藏
Ridged Fields in British Honduras
err1975-01-01
err0
PREAI
errG. W. Olson; A. H. Siemens; D. E. Puleston; G. Cal; D. Jenkins
err分享
err收藏
Progressive Cross-Modal Semantic Network for Zero-Shot Sketch-Based Image Retrieval
err2020-01-01
err55
PREAI
errDeng, Cheng; Xu, Xinxun; Wang, Hao; Yang, Muli; Tao, Dacheng
err分享
err收藏
Early Intra‐Articular Hip Disease Presenting With Posterior Pelvic and Groin Pain
err2009-09-19
err0
PREAI
errHeidi Prather; Devyani Hunt; Anne Fournie; John C. Clohisy
err分享
err收藏
学者 查看更多内容