Return
Looking through the codebook: Generative anomaly segmentation with multi-contrastive learning
DOI:10.1016/j.neucom.2025.132336.png)
Abstract
En 中文
Most existing semantic segmentation models are based on discriminative approaches. Such models often fail to detect out-of-distribution (OoD) objects because they primarily learn class-specific decision boundaries without explicitly modeling the underlying data distribution, which leads to overconfident misclassification of unseen objects. In contrast, generative models aim to capture the underlying data distribution, resulting in more effective anomaly detection. Recent methods still rely on discriminative segmentation networks with generative modules used only as auxiliary components, which prevents them from leveraging generative modeling for pixel-wise likelihood estimation and effective OoD separation. Among generative models, codebook-based approaches such as vector-quantized variational autoencoders (VQ-VAEs) discretize the latent space into a finite set of codevectors, enabling a fully generative formulation in which each pixel is modeled by its likelihood under the learned latent distribution. Motivated by this property, we propose a purely generative anomaly segmentation method that integrates a VQ-VAE, a weighted top- scoring strategy, and multi-contrastive learning. By treating the segmentation class-specific codevectors of VQ-VAE as in-distribution (ID) representations, we introduce a codevector-wise top‑ criterion. This criterion scores each pixel based on its top‑ nearest codevectors and their associated scores, thereby reflecting relative similarities within the same classes. Furthermore, we implement a codevector-based multi-contrastive learning strategy with specially sampled void labels. This implementation effectively structures the latent space among classes and ensures that anomalies are not aligned with ID codevectors. Extensive experiments demonstrate that the proposed method detects anomalies effectively and performs robustly in single- and cross-domain scenarios.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W
Organization
Cited Papers
OA-Pose: Occlusion-aware monocular 6-DoF object pose estimation under geometry alignment for robot manipulation
PATTERN RECOGNITION
IF7.6
CANet: Contextual Information and Spatial Attention Based Network for Detecting Small Defects in Manufacturing Industry
PATTERN RECOGNITION
IF7.6
Deep learning-enhanced environment perception for autonomous driving: MDNet with CSP-DarkNet53
PATTERN RECOGNITION
IF7.6
no more

