arrow
返回

Disentangled image-text classification: Enhancing visual representations with MLLM-driven knowledge transfer

delete2026-01-21
delete0
PRE
AI
Q
Qianjun Shuai
X
Xiaohao Chen *
Y
Yongqiang Cheng
苗
苗方 (Fang Miao)
L
Libiao Jin
DOI:10.1016/j.eswa.2025.130790delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Multimodal image-text classification plays a critical role in applications such as content moderation, news recommendation, and multimedia understanding. Despite recent advances, visual modality faces higher representation learning complexity than textual modality in semantic extraction, which often leads to a semantic gap between visual and textual representations. In addition, conventional fusion strategies introduce cross-modal redundancy, further limiting classification performance. To address these issues, we propose MD-MLLM, a novel image-text classification framework that leverages large multimodal language models (MLLMs) to generate semantically enhanced visual representations. To mitigate redundancy introduced by direct MLLM feature integration, we introduce a hierarchical disentanglement mechanism based on the Hilbert-Schmidt Independence Criterion (HSIC) and orthogonality constraints, which explicitly separates modality-specific and shared representations. Furthermore, a hierarchical fusion strategy combines original unimodal features with disentangled shared semantics, promoting discriminative feature learning and cross-modal complementarity. Extensive experiments on two benchmark datasets, N24News and Food101, show that MD-MLLM achieves consistently stable improvements in classification accuracy and exhibits competitive performance compared with various representative multimodal baselines. The framework also demonstrates good generalization ability and robustness across different multimodal scenarios. The code is available at https://github.com/xiaohaochen0308/MD-MLLM.
Keyword:
Image-text classification
Visual semantic enhancement
Large multimodal language models
Cross-modal representation learning
Disentangled feature fusion

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

U
university of sunderland
学者数:
54
论文数: 44
被引数: 0
C
communication university of china
学者数:
495
论文数: 250
被引数: 0
引用论文

引用论文

Align and Retrieve: Composition and Decomposition Learning in Image Retrieval With Text Feedback
err2024-01-01
err0
PREAI
errXu, Yahui; Bin, Yi; Wei, Jiwei; Yang, Yang; Wang, Guoqing; Shen, Heng Tao
err分享
err收藏
Multimodal Boosting: Addressing Noisy Modalities and Identifying Modality Contribution
err2024-01-01
err1
PREAI
errMai, Sijie; Sun, Ya; Xiong, Aolin; Zeng, Ying; Hu, Haifeng
err分享
err收藏
Multiview adaptive attention pooling for image-text retrieval
err2024-05-01
err3
PREAI
errDing, Yunlai; Yu, Jiaao; Lv, Qingxuan; Zhao, Haoran; Dong, Junyu; Li, Yuezun
err分享
err收藏
Learning From Noisy Correspondence With Tri-Partition for Cross-Modal Matching
err2024-01-01
err0
PREAI
errFeng, Zerun; Zeng, Zhimin; Guo, Caili; Li, Zheng; Hu, Lin
err分享
err收藏
学者 查看更多内容