1
Return

Posterior-calibrated causal circuits in variational autoencoders: why image-domain interpretability fails on tabular data

delete2026-08-01
delete0
PRE
AI
D
Dip Roy *
R
Rajiv Misra
S
Sanjay Kumar Singh
A
Anisha Roy
DOI:10.1007/s00521-026-12322-6delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Mechanistic interpretability has produced substantial insight for discriminative networks, but generative models outside the image domain remain comparatively unexplored. Variational autoencoders (VAEs) are increasingly deployed for tabular imputation, anomaly detection, and synthetic data generation, yet whether mechanistic findings on image-domain VAEs transfer to tabular data has not been tested. We extend a four-level causal intervention framework to four tabular benchmarks (Adult, Credit Default, Bank Marketing, Wine Quality) and one image benchmark (dSprites) across five VAE architectures with three random seeds each, yielding 75 trained models, and add 3DShapes as a second image benchmark (15 supplementary runs) to test whether cross-modality findings generalize. We introduce three methodological refinements (posterior-calibrated Causal Effect Strength (CES), path-specific activation patching, and Feature-Group Disentanglement) and report all CES architecture comparisons before and after Frisch-Waugh residualization on reconstruction MSE to address whether CES reflects circuit structure or decoder competence. We find that tabular VAE circuits exhibit approximately 33 percent lower modularity than synthetic image benchmarks (image-to-tabular ratio 1.49 across pooled image runs), β-VAE shows substantially weaker per-dimension causal influence on tabular data than on image data (tabular pooled CES = 0.043 vs pooled image CES = 0.107) with a 260 × reduction on Adult Income relative to Standard VAE and phase-transition behaviour between β = 2.0 and β = 4.0, and most of the CES architectural signal is reconstruction-mediated (only 3 of 9 originally-significant pairwise comparisons survive MSE residualization, all involving DIP-VAE-II). Specificity emerges as the most reconstruction-independent discriminative metric (raw and partial correlations with downstream AUC differ by less than 0.012, r = 0.460, p < 0.001), and imputation under random feature missingness reveals a task-dependent reversal in which Specificity is anti-predictive (partial r = + 0.702, p < 0.001) while CES becomes positively predictive (partial r = -0.453, p = 0.0003). Architectural guidance derived from image-domain VAE studies does not transfer to tabular data, and per-architecture rankings differ even across synthetic image benchmarks. Practitioners should report reconstruction MSE alongside any circuit metric, prefer Specificity over Modularity for classification-oriented tabular VAEs, and validate β empirically before deploying β-VAE on tabular data with substantial feature redundancy.
Keywords:
Mechanistic interpretability
Variational autoencoders
Cross-modality transfer
Causal interventions
Tabular data
Disentanglement
Circuit analysis
Causal effect strength

Journal

Neural Computing and Applications cover
Neural Computing and Applications
IF:
4.5
Papers:
729
Citations:
3.2W

Organization

D
department of computer science
Scholars:
547
Papers: 287
Citations: 0
D
department of computer science and engineering
Scholars:
1.7K
Papers: 961
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers