Return
A causality based multi-task framework for enhancing image and text interactions
Z
Z
Y
DOI:10.1016/j.knosys.2026.116796.png)
Abstract
En 中文
• It addresses overlooked, inconsistent image–text interactions in common tasks. • The framework further improves the model’s performance, even if the model is SoTA. • The causality aims to explore more crucial features instead of finding the noise. • The modeling is end-to-end, thus avoiding computational overhead in similar works.
Keywords:
Multi-modal interaction
Multi-task modeling
Image and text recognition
Causality intervention
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W
