1
Return

A causality based multi-task framework for enhancing image and text interactions

delete2026-08-08
delete0
PRE
AI
Z
Zhaomeng Cheng *
Z
Zhong Ji *
Y
Yan Zhang
DOI:10.1016/j.knosys.2026.116796delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• It addresses overlooked, inconsistent image–text interactions in common tasks. • The framework further improves the model’s performance, even if the model is SoTA. • The causality aims to explore more crucial features instead of finding the noise. • The modeling is end-to-end, thus avoiding computational overhead in similar works.
Keywords:
Multi-modal interaction
Multi-task modeling
Image and text recognition
Causality intervention

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

T
tianjin university
Scholars:
7.7W
Papers: 5.6W
Citations: 88
T
tsinghua university
Scholars:
11.5W
Papers: 9.9W
Citations: 137
Cited Papers

Cited Papers

Citing Papers

Citing Papers