arrow
Return

Alignment-Invertibility Regularization for Explainable Neural Networks

delete2026-02-17
delete0
PRE
AI
B
Borui Zhang
Q
Qihang Rao
周杰 (Jie Zhou)
J
Jiwen Lu
DOI:10.1109/TPAMI.2026.3665728delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep learning has profoundly impacted society, yet the inherent nature of deep neural networks hinders further application to high-reliability industries. To demystify these closed-boxes, numerous works attempt to improve the explainability by observing or impacting internal variables of the models. However, existing methods rely on heuristics without rigorous theoretical foundations, often requiring intricate model modifications or redesigns. This work first formalizes two fundamental properties of explainability: <b>alignment</b> and <b>invertibility</b>, serving as theoretical pillars for rigorous interpretability analysis. Building on these, we introduce <b>Bort</b>, a plug-and-play optimizer that enforces <u><b>B</b></u>oundedness and <u><b>ort</b></u>hogonality constraints on model parameters to improve explainability. These constraints are theoretically derived from the alignment and invertibility principles. Considering conventional optimizers can not leverage data features for precise attribution, we present a data-aware extension, termed <b>DBort</b>, which integrates an auxiliary loss term. Intriguingly, in the linear case, DBort converges to Principal Component Analysis (PCA). Our in-depth analysis of penalty term design reveals that <inline-formula><tex-math notation="LaTeX">$l_{1}$</tex-math></inline-formula>-based penalties provide a more stringent adherence to the imposed constraints compared to their <inline-formula><tex-math notation="LaTeX">$l_{2}$</tex-math></inline-formula> counterparts. Our experiments involve reconstructing and backtracking through the optimized model representations, which reveal a marked enhancement in explainability. Furthermore, leveraging Bort, we successfully synthesize explainable adversarial examples without additional training. Notably, Bort consistently improves the classification accuracy across diverse architectures, including ResNet and DeiT, on benchmark datasets such as MNIST, CIFAR-10, and ImageNet.
Keywords:
Interpretable deep learning
XAI
neural networks
optimizer
explainability

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

T
tsinghua university
Scholars:
11.7W
Papers: 9.9W
Citations: 137