arrow
Return

Multimodal interpretable image recognition network via language-guided global-local collaboratively alignment

delete2026-01-28
delete0
PRE
AI
张素兰 (Sulan Zhang)
P
Peijun Zhang
胡立华 (Lihua Hu)
X
Xin Wen
张继福 (Jifu Zhang)
DOI:10.1016/j.knosys.2026.115422delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• A language-guided global–local collaborative alignment framework is proposed for multimodal interpretable image recognition. • Dual filtering strategies are designed to perform concept de-redundancy and visual recognizability verification to ensure high-quality textual concepts generated by large language models. • A local visual prompt module is introduced to identify fine-grained visual regions for precise local vision–language alignment. • A learnable dynamic weighting mechanism is developed to adaptively fuse global and local semantic information to enhance accuracy and interpretability. • Experiments on both general and fine-grained datasets demonstrate that LGLCA-Net not only provides finer-grained and more trustworthy visual explanations but also improves recognition performance.

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

T
taiyuan university of science and technology
Scholars:
1.1K
Papers: 420
Citations: 0