Return
Enhance multi-modal structured representations with open information extraction
DOI:10.1016/j.engappai.2025.112903.png)
Abstract
En 中文
Current large-scale vision-language models are expanding their range of applications and have achieved impressive performance in multi-modal tasks. However, the performance of existing models in expressing structured semantic information is of concern, as they have difficulty distinguishing the relationship between subjects and objects in some scene-specific images. This is because the feature learning process in multi-modal task scenarios does not incorporate structured knowledge into the model. In this study, we propose an enhanced end-to-end contrastive language-image pre-training (CLIP) model with open information extraction (OIE-CLIP), which is used to assist in the training of multi-modal models for short texts by integrating structured knowledge representations to enhance the ability of multi-modal representation of structured information. OIE-CLIP leverages the construction of effective negative examples to enhance contrastive learning. In addition, we propose a triple knowledge encoder (TKE) based on the output of open information extraction (OIE) to further boost the structured representation capability of the multi-modal model. To verify the validity of our approach, we pre-trained the model with the above method and downstream experiments in multi-modal tasks. The experimental results show that our method performs best on the public datasets visual genome attribution (VG-Attribution) and visual genome relation (VG-Relation), outperforming the multi-modal state-of-the-art model by 2.2% and 1.8%, respectively. Furthermore, our experimental results on the Microsoft Common Objects in Context (MSCOCO) dataset further demonstrate that our method can effectively improve the structured representations.
Journal
IF:
8
Papers:
5.4K
Citations:
3.5W
Organization
No organization information available

