1
Return

A unifying framework from neural superposition to sparse interpretable codes

delete2026-07-14
delete0
PRE
AI
D
David Klindt *
C
Charles O’Neill
P
Patrik Reizinger
H
Harald Maurer
N
Nina Miolane
DOI:10.1038/s42256-026-01259-zdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Understanding how information is represented in neural networks is a fundamental challenge in both neuroscience and artificial intelligence. Despite their nonlinear computations, ample evidence suggests that neural networks encode features in superposition, meaning that systems linearly represent more concepts than they have neurons. This observation opens the door to extracting interpretable representations from otherwise opaque networks, but a principled account of why superposition arises and how it can be exploited has been lacking. Here we synthesize insights from identifiability theory, compressed sensing and quantitative interpretability research to propose a unified perspective on this phenomenon. Our synthesis yields a three-step framework: identifiability theory establishes that neural networks trained for classification recover latent features up to linear mixing, compressed sensing provides guarantees for disentangling these features via sparse coding, and interpretability metrics grounded in behavioural tasks assess whether the extracted features align with human-interpretable concepts. By bridging theoretical neuroscience, representation learning and interpretability research, our framework connects longstanding questions about neural coding in biological systems with modern efforts in artificial intelligence transparency, and highlights open problems at their intersection. Kindt et al. present a unifying framework for superposition in neural networks. Their three-step approach clarifies how latent features can be identified, disentangled and assessed.

Journal

Nature Machine Intelligence cover
Nature Machine Intelligence
IF:
23.9
Papers:
1.3K
Citations:
1.5W

Organization

C
Cold Spring Harbor Laboratory
Scholars:
2.8K
Papers: 1.7K
Citations: 7.2K
A
australian national university
Scholars:
1.9K
Papers: 1.0K
Citations: 0
E
ellis institute
Scholars:
2
Papers: 1
Citations: 0
U
University of California Santa Barbara
Scholars:
1.2W
Papers: 9.5K
Citations: 3.6W
U
University of Tübingen
Scholars:
295
Papers: 122
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers