arrow
Return

Exploring refined dual visual features cross-combination for image captioning

delete2024-12-01
delete0
PRE
AI
J
Junbo Hu
Z
Zhixin Li *
Q
Qiang Su
Z
Zhenjun Tang
马慧芳 (Huifang Ma)
DOI:10.1016/j.neunet.2024.106710delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
For current image caption tasks used to encode region features and grid features Transformer-based encoders have become commonplace, because of their multi-head self-attention mechanism, the encoder can better capture the relationship between different regions in the image and contextual information. However, stacking Transformer blocks necessitates quadratic computation through self-attention to visual features, not only resulting in the computation of numerous redundant features but also significantly increasing computational overhead. This paper presents a novel Distilled Cross-Combination Transformer (DCCT) network. Technically, we first introduce a distillation cascade fusion encoder (DCFE), where a probabilistic sparse self-attention layer is used to filter out some redundant and distracting features that affect attention focus, aiming to obtain more refined visual features and enhance encoding efficiency. Next, we develop a parallel cross-fusion attention module (PCFA) that fully exploits the complementarity and correlation between grid and region features to better fuse the encoded dual visual features. Extensive experiments conducted on the MSCOCO dataset demonstrate that our proposed DCCT method achieves outstanding performance, rivaling current state-of-the-art approaches.
Keywords:
Image captioning
Cross Combination
Contrastive Language-Image Pre-Training
Reinforcement learning

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.8K
Citations:
3.0W

Organization

G
Guangxi Normal University
Scholars:
7.7K
Papers: 4.9K
Citations: 5.1K
N
northwest normal university - china
Scholars:
7.8K
Papers: 4.8K
Citations: 4