arrow
Return

Unsupervised Background Subtraction Using Generator-Discriminator Learning

delete2025-08-12
delete0
PRE
AI
B
Basit Alawode
S
Sajid Javed
DOI:10.1109/TCSVT.2025.3598078delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Background subtraction is a core problem in computer vision, widely used in video surveillance to segment moving foreground objects from video sequences. While deep learning approaches have shown strong performance—especially under dynamic backgrounds and sudden illumination changes—they typically rely on large-scale, high-quality labeled video datasets. Acquiring such data is time-consuming and expensive, making existing supervised or weakly supervised methods less suitable for real-time applications. Moreover, many of these methods suffer from performance degradation when applied to unseen video sequences. To address these challenges, we present UTGMP-BS algorithm: an Unsupervised Transformer-based pseudo-label Generator with a Message-Passing network for the Background Subtraction task. UTGMP-BS is a fully unsupervised framework designed to learn directly from unlabeled video sequences. It comprises two key components: a transformer-based pseudo-label generator, which produces initial pixel-level foreground and background labels using an encoder-decoder architecture and an $\mathcal{L}_{1}$ loss, and a message-passing network, which acts as a label-cleaner discriminator to refine the pseudo labels and enforce spatial consistency. These two branches engage in mutual learning through consecutive iterations, enhancing one another’s performance without any ground-truth supervision. The framework is trained using an alternating iterative learning strategy with binary cross-entropy loss, achieving robust background subtraction across varied scenes. Extensive experiments on six publicly available benchmark datasets demonstrate that UTGMP-BS achieves competitive results compared to existing State-of-The-Art (SOTA) methods.
Keywords:
Background subtraction
moving object segmentation
background modeling
unsupervised learning
vision transformer
message-passing network

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
612
Citations:
3.1W

Organization

K
khalifa university of science and technology
Scholars:
76
Papers: 36
Citations: 0