Return
Multimodal interpretable image recognition network via language-guided global-local collaboratively alignment
DOI:10.1016/j.knosys.2026.115422.png)
Abstract
En 中文
• A language-guided global–local collaborative alignment framework is proposed for multimodal interpretable image recognition. • Dual filtering strategies are designed to perform concept de-redundancy and visual recognizability verification to ensure high-quality textual concepts generated by large language models. • A local visual prompt module is introduced to identify fine-grained visual regions for precise local vision–language alignment. • A learnable dynamic weighting mechanism is developed to adaptively fuse global and local semantic information to enhance accuracy and interpretability. • Experiments on both general and fine-grained datasets demonstrate that LGLCA-Net not only provides finer-grained and more trustworthy visual explanations but also improves recognition performance.
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W

