arrow
Return

Diffusion-Driven RGB-D Salient Object Detection With Temporal Modulation

delete2026-04-17
delete0
PRE
AI
S
Shixiang Shi
G
Gongyang Li
R
Runmin Cong
肖顺鑫 cover
肖顺鑫 (Shunxin Xiao)
W
Weisi Lin
DOI:10.1109/tcsvt.2026.3684803delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Existing RGB-D Salient Object Detection (SOD) methods are primarily built on the end-to-end prediction paradigm. Although these methods have achieved remarkable progress, they still struggle to generate accurate predictions in some complex scenes due to their lack of error correction capability. In this paper, we explore the use of conditional diffusion architectures for RGB-D SOD, producing saliency maps in a step-by-step generation paradigm. Accordingly, we propose <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DiffRGBD</i>, a novel diffusion-driven framework with temporal modulation. The core of DiffRGBD is using time steps to control the conditional information injected into the denoising network in a two-stage temporal modulation manner. Specifically, our DiffRGBD comprises a feature extractor, a conditional generator, two temporal modulators, and a denoising network. First, the SAM2 encoder with adapters is adopted to extract hierarchical cross-modal features. Then, the Mutual-Differential Attention Module is responsible for generating the conditional information via effective cross-modal fusion. Notably, the conditional information continuously achieves channel modulation and spatial modulation in the Temporal Channel Enhancement Module and the Temporal Spatial Refinement Module (<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i</i>.<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">e</i>., two temporal modulators), resulting in comprehensive conditional information. Finally, conditional information is injected into the denoising network to guide the production of saliency maps. As the time step increases, our DiffRGBD can gradually correct errors and generate accurate saliency maps. Extensive experiments on seven public RGB-D SOD benchmarks demonstrate that our proposed DiffRGBD achieves superior performance over state-of-the-art methods. The code and results of our method are available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/Shixiang02/DiffRGBD</uri>
Keywords:
RGB-D salient object detection
conditional diffusion model
temporal modulation
cross-modal fusion

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
612
Citations:
3.1W

Organization

N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.8W
Citations: 8.1W
S
shandong university
Scholars:
9.3W
Papers: 6.4W
Citations: 94
X
Xiamen University of Technology
Scholars:
3.7K
Papers: 2.5K
Citations: 5.1K
S
shanghai university
Scholars:
3.8W
Papers: 2.7W
Citations: 52
researcher View more organizations