返回
Disentangled image-text classification: Enhancing visual representations with MLLM-driven knowledge transfer
DOI:10.1016/j.eswa.2025.130790.png)
摘要
En 中文
Multimodal image-text classification plays a critical role in applications such as content moderation, news recommendation, and multimedia understanding. Despite recent advances, visual modality faces higher representation learning complexity than textual modality in semantic extraction, which often leads to a semantic gap between visual and textual representations. In addition, conventional fusion strategies introduce cross-modal redundancy, further limiting classification performance. To address these issues, we propose MD-MLLM, a novel image-text classification framework that leverages large multimodal language models (MLLMs) to generate semantically enhanced visual representations. To mitigate redundancy introduced by direct MLLM feature integration, we introduce a hierarchical disentanglement mechanism based on the Hilbert-Schmidt Independence Criterion (HSIC) and orthogonality constraints, which explicitly separates modality-specific and shared representations. Furthermore, a hierarchical fusion strategy combines original unimodal features with disentangled shared semantics, promoting discriminative feature learning and cross-modal complementarity. Extensive experiments on two benchmark datasets, N24News and Food101, show that MD-MLLM achieves consistently stable improvements in classification accuracy and exhibits competitive performance compared with various representative multimodal baselines. The framework also demonstrates good generalization ability and robustness across different multimodal scenarios. The code is available at https://github.com/xiaohaochen0308/MD-MLLM.
Keyword:
Image-text classification
Visual semantic enhancement
Large multimodal language models
Cross-modal representation learning
Disentangled feature fusion
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
引用论文
A Transformer-Based Model With Self-Distillation for Multimodal Emotion Recognition in Conversations基于Transformer的对话中多模态情感识别的自蒸馏模型
Factual consistency evaluation of summarization in the Era of large language models大语言模型时代摘要的事实一致性评价
Cross-modal contrastive learning for multimodal sentiment recognition面向多模态情感识别的跨模态对比学习
APPLIED INTELLIGENCE
IF3.5
Weakening the Dominant Role of Text: CMOSI Dataset and Multimodal Semantic Enhancement Network弱化文本的主导作用: CMOSI数据集和多模态语义增强网络

