Return
DDNet: A Dual-Optimized Diffusion Framework for Camouflaged Object Detection in IoT-Enabled Visual Surveillance
H
C
S
DOI:10.1109/jiot.2026.3704296.png)
Abstract
En 中文
Camouflaged object detection (COD) in the Internet-of-Things visual sensing scenarios remains challenging due to the extreme feature homogeneity between targets and backgrounds. Conventional semantic segmentation approaches suffer from overconfident predictions and information redundancy, leading to degraded boundary delineation under complex camouflage conditions. We propose DDNet, a dual-optimization framework that synergistically integrates discriminative feature extraction with conditional diffusion-based mask generation. The framework comprises two complementary modules: a global feature extraction module (GFEM) that employs multiscale convolutions with channel attention gating to suppress spatial redundancy while preserving semantic context, and a local feature extraction module (LFEM) that leverages bidirectional adaptive pooling and hierarchical dilated convolutions to enhance fine-grained boundary representations. These features condition a denoising diffusion probabilistic model (DDPM), which performs iterative noise injection and progressive refinement to generate high-fidelity segmentation masks through latent space optimization. A signal-to-noise ratio (SNR)-informed schedule is adopted to mitigate mask structure degradation during the diffusion process. Extensive experiments on three benchmark datasets (CAMO, COD10K, and NC4K) demonstrate that DDNet achieves state-of-the-art performance, outperforming 16 existing methods across four evaluation metrics. Furthermore, cross domain evaluations on polyp segmentation confirm the generalization capability of the proposed framework, underscoring its potential for diverse IoT-enabled perception tasks.
Keywords:
Camouflaged object detection (COD)
deep neural network
diffusion model
semantic segmentation
Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W
