1
Return

Boosting Multimodal Chain of Thought Reasoning by Selective Mixture of Experts

delete2026-08-06
delete0
delete
OA
AI
Q
Qilei Li
S
Shitong Sun
D
Da Li *
T
Timothy Hospedales
S
Shaogang Gong
DOI:10.1016/j.patcog.2026.114570delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
• Multimodal CoT lacks visual knowledge discovery for robust reasoning. • Introduce SESA, integrating visual and textual features for reasoning. • SESA links visual encoders (e.g., ViT) with LMs (e.g., Flan-T5) for fine-tuning. • Easily integrates with models such as MM-CoT and DD-CoT.
Keywords:
Multimodal reasoning
Chain of Thought
Mixture of Experts
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

Q
Queen Mary University of London
Scholars:
945
Papers: 466
Citations: 3.4W
U
University of York
Scholars:
1.1K
Papers: 631
Citations: 2.5W
C
central china normal university
Scholars:
2.3K
Papers: 856
Citations: 0
U
University of Edinburgh
Scholars:
5.1W
Papers: 4.5W
Citations: 70
Cited Papers

Cited Papers

Citing Papers

Citing Papers