Return
Boosting Multimodal Chain of Thought Reasoning by Selective Mixture of Experts
Q
S
D
T
S
DOI:10.1016/j.patcog.2026.114570.png)
Abstract
En 中文
• Multimodal CoT lacks visual knowledge discovery for robust reasoning. • Introduce SESA, integrating visual and textual features for reasoning. • SESA links visual encoders (e.g., ViT) with LMs (e.g., Flan-T5) for fine-tuning. • Easily integrates with models such as MM-CoT and DD-CoT.
Keywords:
Multimodal reasoning
Chain of Thought
Mixture of Experts
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W
