Return
Key features-guided Multi-View Collaborative Network for image captioning
DOI:10.1016/j.neunet.2025.107814.png)
Abstract
En 中文
The integration of multi-view has significantly advanced the image captioning task. However, the semantic noise introduced during this integration process presents a challenge, constraining further performance improvement. To overcome this challenge, we propose a Key features-guided Multi-view Collaborative Network (KMCN), a novel approach that achieves multi-view complementary advantages and minimizes semantic noise during critical sentence prediction steps, including feature enhancement and cross-modal semantic alignment. Specifically, we introduce a Key features-guided Augmentation and Fusion Encoder (KAFE) in the feature enhancement stage, which adopts key features to provide the necessary complementary information. It is conducive to generating refined multi-view feature representation while reducing potential semantic noise introduced by meaningless interactions. Subsequently, we introduce a Dual-branch Collaborative Decoder (DCD) in the cross-modal semantic alignment stage, modeling inter-modal relationships through cross-guided dual-branch block. The design aims to decompose complex multi-view feature space into multiple relatively simple subspaces during the decoding process. It is conducive to avoiding semantic noise caused by distorted mappings. To validate the performance of KMCN, we conduct extensive experiments on the highly competitive Microsoft Common Objects in Context (MS-COCO) benchmark dataset. The results show that our KMCN outperforms several state-of-the-art image captioning models on both offline and online tests.
Journal
IF:
6.3
Papers:
7.8K
Citations:
3.0W
Organization
No organization information available

