Return
STAFuse: Scene-Text Aggregation Guided Composite Degradation-Robust Infrared and Visible Image Fusion
T
H
DOI:10.1109/tip.2026.3719466.png)
Abstract
En 中文
Infrared and visible image fusion aims to integrate complementary information from source images to generate high-quality fusion images that serve downstream tasks. However, the differentiated representation of image scene content, the unpredictability of degradation modes in source images, and the complexity of composite degradations pose significant challenges to constructing degradation-robust fusion models. To overcome these challenges, this study proposes STAFuse, a framework designed for composite degradation-robust image fusion that utilizes adaptive degradation-mode identification and aggregated textual prior guidance. We first introduce a scene- and degradation-aware mechanism that extracts crucial context and degradation data, converting it into a textual format to generate dynamic convolution kernels. This allows for the adaptive identification and elimination of varying degradation effects and scene disparities. Additionally, we implement an aggregated scene prior that condenses multi-source information into a fusion text, simulating an ideal scene to effectively guide the retrieval and fusion of multimodal information. Finally, we design a textual-domain supervision loss to perform auxiliary semantic supervision on the fusion model, thereby suppressing degradation effects and improving the visual fidelity of the fusion results. Experimental results demonstrate that the proposed method significantly outperforms existing state-of-the-art methods in terms of robustness, flexibility, and the aggregation of complementary information.
Keywords:
Image Fusion
composite degradation
textual guidance
CondConv
Journal
IF:
13.7
Papers:
1.0W
Citations:
8.4W
