arrow
Return

Self-Supervised Disentangled Representation Learning via Compositional Invariance

delete2025-10-23
delete0
PRE
AI
H
Haoqiang Chen
J
Jianxiang Sun
Y
Yadong Liu
H
Hao Shi
D
Dewen Hu
DOI:10.1109/TCSVT.2025.3625076delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Images serve as a crucial information source for machine intelligence to understand the world, while how to represent images significantly impacting the generalizability and interpretability of intelligent systems. Disentangled representation learning offers a promising approach to improve both aspects. However, most of existing methods predominantly rely on statistical independence assumptions. This poses two key limitations: first, it fails to capture the reality that many concepts are both disentangled yet interrelated; second, it conflicts with human cognitive patterns where concepts naturally exhibit complex dependencies. These limitations further hinder collaboration between machine and human beings. To overcome these limitations, we propose Compositional Invariant Disentanglement (CID), a novel self-supervised learning method that enables models to learn composable representations aligned with human cognitive habits. Inspired by humans’ ability to flexibly recombine concepts, we reframe the definition of disentanglement through the lens of compositional invariance rather than statistical independence. This paradigm shift allows effective disentanglement even with correlated factors, achieving state-of-the-art disentanglement performance across multiple standard benchmarks (improved by 4.0% on Shapes3D, 4.5% on Dsprites, and 28.6% on MPI3D). Furthermore, by building upon and extending the successful self-supervised learning framework BYOL, CID demonstrates potential for large-scale disentanglement pre-training on unlabeled data. This work contributes to extracting more robust and interpretable representations from images for machine intelligence.
Keywords:
Disentangled representation learning
compositional invariance
self-supervised learning
generalization
interpretability

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
612
Citations:
3.1W

Organization

N
national university of defense technology
Scholars:
4.4K
Papers: 1.4K
Citations: 0