Return
Self-Supervised Disentangled Representation Learning via Compositional Invariance
DOI:10.1109/TCSVT.2025.3625076.png)
Abstract
En 中文
Images serve as a crucial information source for machine intelligence to understand the world, while how to represent images significantly impacting the generalizability and interpretability of intelligent systems. Disentangled representation learning offers a promising approach to improve both aspects. However, most of existing methods predominantly rely on statistical independence assumptions. This poses two key limitations: first, it fails to capture the reality that many concepts are both disentangled yet interrelated; second, it conflicts with human cognitive patterns where concepts naturally exhibit complex dependencies. These limitations further hinder collaboration between machine and human beings. To overcome these limitations, we propose Compositional Invariant Disentanglement (CID), a novel self-supervised learning method that enables models to learn composable representations aligned with human cognitive habits. Inspired by humans’ ability to flexibly recombine concepts, we reframe the definition of disentanglement through the lens of compositional invariance rather than statistical independence. This paradigm shift allows effective disentanglement even with correlated factors, achieving state-of-the-art disentanglement performance across multiple standard benchmarks (improved by 4.0% on Shapes3D, 4.5% on Dsprites, and 28.6% on MPI3D). Furthermore, by building upon and extending the successful self-supervised learning framework BYOL, CID demonstrates potential for large-scale disentanglement pre-training on unlabeled data. This work contributes to extracting more robust and interpretable representations from images for machine intelligence.
Keywords:
Disentangled representation learning
compositional invariance
self-supervised learning
generalization
interpretability
Journal
IF:
11.1
Papers:
612
Citations:
3.1W

