Return
Target-oriented consistent cross-modal alignment framework for multimodal stance detection
DOI:10.1016/j.neucom.2026.133398.png)
Abstract
En 中文
• We propose the TCCA method, which leverages the stance target to guide the learning of textual and visual representations. This design simplifies the multimodal alignment process and enhances the model’s ability to capture fine-grained stance cues. • We design the Target-oriented Prompt as Visual-Language Query module to activate target-relevant visual and textual representations, enabling the model to focus on stance-critical information. Building on this, we construct the Target-oriented Cross-modal Alignment module to bring target-oriented textual and visual semantics closer together, fostering a unified and cohesive semantic space. • Extensive experiments on the MMSD and MMVax-STANCE datasets demonstrate that our method achieves superior performance over existing approaches in most evaluation metrics.
Keywords:
TCCA
multimodal stance detection
target-oriented prompt
cross-modal alignment
visual-language query
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

