Return
Toward Complex Backgrounds: A Unified Difference-Aware Decoder for Binary Segmentation
DOI:10.1109/TCSVT.2025.3612574.png)
Abstract
En 中文
Binary segmentation is used to distinguish objects of interest from background, and is an active area of convolutional encoder-decoder network research. The current decoders are designed for specific objects based on the common backbones as the encoders, but cannot deal with complex backgrounds. Inspired by the way human eyes detect objects, we propose a new unified dual-branch decoder paradigm, termed the difference-aware decoder, to better explore the differences between foreground and background and to separate objects of interest in optical images. This decoder operates in two stages, leveraging multi-level features from the encoder. In the first stage, coarse detection of foreground objects is achieved by directly utilizing high-level semantic features, mimicking the initial rough observation of human vision. In the second stage, the decoder refines segmentation by exploring differences in low-level features, guided by the coarse map from the first stage. To enhance this process, we introduce two key innovations. First, a difference-aware prototype generation strategy leverages the guide map to extract foreground and background prototypes from high-level features, and calculates the similarity between these prototypes and corresponding representations in low-level feature spaces. Second, an overlapped window cross-level semantic guidance mechanism integrates high-level semantic information into low-level features through channel grouping and multi-scale aligned window pairs, guided by the similarities computed in the first strategy. Together, these innovations significantly enhance the DAD’s ability to discern subtle differences, enabling precise foreground extraction and effectively addressing the challenges of complex and varied backgrounds. To verify the performance of the proposed difference-aware decoder, we choose three well known backbones including ResNet, Res2Net, PVT, and two binary segmentation tasks, <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$i.e$ </tex-math></inline-formula>., salient object detection, and camouflaged object detection, for comparative experiments. The results demonstrate that the difference-aware decoder can achieve higher accuracy than the other state-of-the-art binary segmentation methods for these tasks. The source code will be available on <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/Henryjiepanli/DAD</uri>
Keywords:
Dual-branch decoder
binary segmentation
salient object detection
camouflaged object detection
polyp segmentation
Journal
IF:
11.1
Papers:
612
Citations:
3.1W

