Return
MSSCR: Multi-scale semantic collaborative reasoning model for explainable multimodal rumor detection
DOI:10.1016/j.neucom.2025.131042.png)
Abstract
En 中文
Multimodal rumors, which leverage visual “evidence” to enhance their credibility, are more deceptive than textual rumors on online social media. This makes multimodal rumor detection a critical area of research in the field of rumor analysis. While existing methods primarily rely on the coarse-grained alignment between images and text, the increasing sophistication of image generation and manipulation technologies has rendered such approaches inadequate for precise rumor detection. To address this challenge, we propose a Multi-Scale Semantic Collaborative Reasoning model (MSSCR) for explainable multimodal rumor detection. Our model explores complementary relationships from macro, meso, and micro perspectives. At the macro level, a co-attention mechanism is employed to align and fuse cross-modal global semantics, capturing holistic relationships among the post, its comments, and the associated image. At the meso level, we construct graph structures to model local interactions between the post and its comments, as well as between the image and the comments. The Graph Attention Network (GAT) is utilized to extract structured local semantic features. At the micro level, a Large Language Model (LLM) is used to extract fine-grained semantics, including keyword and sentiment consistency, to identify subtle semantic discrepancies. Finally, a dynamic weighting mechanism integrates multi-scale features, and an additional fine-grained semantic consistency score derived from micro-level semantics further refines the final rumor prediction. Experimental results on two real-world datasets, PHEME and Weibo, demonstrate that our model achieves improvements of 1.08 % and 0.82 % over the best baseline method, respectively.
Keywords:
multimodal rumors
rumor detection
multi-scale semantic reasoning
graph attention network
large language model

