Return
Stochastic PAC-Bayesian transformers for network intrusion detection and natural language processing applications
DOI:10.1007/s10115-026-02824-z.png)
Abstract
En 中文
Transformer architectures dominate contemporary machine learning, yet face critical limitations in security applications: vulnerability to adversarial attacks, lack of calibrated uncertainty estimates, and difficulty distinguishing confident predictions from uncertain cases requiring human review. We introduce stochastic probably approximately correct (PAC) Bayesian transformers that convert deterministic attention into probabilistic variants via variational inference, unifying uncertainty quantification, and adversarial robustness within a single framework. Our approach replaces fixed attention weights with learned variational distributions and propagates uncertainty through Monte Carlo (MC) sampling, creating moving targets that degrade adversarial effectiveness. We derive joint PAC-Bayesian bounds showing that parameter stochasticity improves both calibration and robustness, with complexity scaling as $$O\left( {\sqrt {KL\left( {\rho ||\pi } \right)/n} } \right)$$ , where KL denotes the Kullback–Leibler divergence between the learned posterior $$\rho$$ and prior $$\pi$$ , and $$n$$ is the sample size. Across network intrusion detection, toxic content detection, and fake news identification, we achieve $$96.8 \pm 0.8\%$$ accuracy with the expected calibration error (ECE) of $$0.043 \pm 0.006$$ , and maintain $$88.3 \pm 1.5\%$$ robust accuracy under multiple attack strategies. Active learning guided by uncertainty reduces labeling requirements by $$68{\text{\% }}$$ , reaching $$95{\text{\% }}$$ of full-data performance with only $$35{\text{\% }}$$ of labels.
Keywords:
Adversarial machine learning
Transformer architectures
Bayesian neural networks
Uncertainty quantification
Adversarial robustness
PAC-Bayesian theory
Journal
IF:
3.1
Papers:
530
Citations:
5.2K

