arrow
Return

Multi-Perspective Cross-Modal Object Encoding for Referring Expression Comprehension

delete2025-01-01
delete0
PRE
AI
J
Jingcheng Ke
文杰 cover
文杰 (Jie Wen)
H
Hui‐Ting Wang
W
Wen-Huang Cheng
J
Jia Wang
DOI:10.1109/TIP.2025.3620129delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Referring expression comprehension (REC) is a crucial task in understanding how a given text description identifies a target object within an image. Existing two-stage REC methods have demonstrated strong performance due to their rational framework design. However, during the encoding of object candidates in an image, most two-stage methods rely exclusively on features extracted from pre-trained detectors, often neglecting the contextual relationships between an object and its neighboring elements. This limitation hinders the full capture of contextual and relational information, reducing the discriminative power of object representations and negatively impacting subsequent processing. In this paper, we propose two novel plug-and-adapt modules: expression-guided label representation module (ELR) and cross-modal calibrated semantic module (CCS), designed to enhance two-stage REC methods. Specifically, the ELR module connects the noun phases of expression to the categorical labels of object candidates in the image, ensuring effective alignment between them. Guided by these connections, a CCS module is introduced to represent each object candidate by integrating its features with those of neighboring candidates from multiple perspectives. This preserves the intrinsic information of each candidate while incorporating relational cues from other objects, enabling more precise embeddings and effective downstream processing in two-stage REC methods. Extensive experiments on six datasets demonstrate the importance of incorporating prior statistical knowledge, and detailed analysis shows that the proposed modules strengthen the alignment between image and text. As a result, our method achieves competitive performance and is compatible with most two-stage methods in the REC task. The code is available on Github: https://github.com/freedom6927/ELR_CCS.git.
Keywords:
Referring expression comprehension
expression-guided label representation module
cross-modal calibrated semantic module

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

G
Guangdong Pharmaceutical University
Scholars:
7.9K
Papers: 3.8K
Citations: 5.3K
H
harbin institute of technology
Scholars:
8.0W
Papers: 6.6W
Citations: 66
N
National Taiwan University
Scholars:
4.7W
Papers: 4.2W
Citations: 3.6W
O
osaka university
Scholars:
2.6W
Papers: 1.9W
Citations: 30
G
Guangdong Polytechnic Normal University
Scholars:
1.6K
Papers: 1.4K
Citations: 1.1K
researcher View more organizations