arrow
返回

PRNet: A Progressive Refinement Network for referring image segmentation

delete2025-05-01
delete0
PRE
AI
J
Jing Liu
H
Huajie Jiang *
Y
Yongli Hu
B
Baocai Yin
DOI:10.1016/j.neucom.2025.129698delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The effective feature alignment between language and image is necessary for correctly inferring the location of reference instances in the referring image segmentation (RIS) task. Previous studies usually resort to assisting target localization with the help of external detectors or using a coarse-grained positional prior during multimodal feature fusion to implicitly enhance the modal alignment capability. However, these approaches are either limited by the performance of the external detector and the design of the matching algorithm, or ignore the fine-grained features in the reference information when using the coarse-grained prior processing, which may lead to inaccurate segmentation results. In this paper, we propose anew RIS network, Progressive Refinement Network (PRNet), which aims to gradually improve the alignment quality between language and image from coarse to fine. The core of the PRNet is the Progressive Refinement Localization Scheme (PRLS), which consists of a Coarse Positional Prior Module (CPPM) and a Refined Localization Module (RLM). The CPPM obtains rough prior positional information and corresponding semantic features by calculating the similarity matrix between sentence and image. The RLM fuses information from the visual and language modalities by densely aligning pixels with word features and utilizes the prior positional information generated by the CPPM to enhance the textual semantic understanding, thus guiding the model to perceive the position of the reference instance more accurately. Experimental results show that the proposed PRNet performs well on all three public datasets, RefCOCO, RefCOCO+, and RefCOCOg.
Keyword:
Referring image segmentation
Position prior
Features alignment
Progressive localization
Transformer

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

B
Beijing University of Technology
学者数:
2.8W
论文数: 2.1W
被引数: 2.7W
引用论文

引用论文

Object detection and recognition via clustered features
err2018-12-01
err37
errOAAI
errWozniak, Marcin; Polap, Dawid
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
学者 查看更多内容