Return
A Provable Optimization Framework for Knowledge Forgetting in Transformer Models
DOI:10.1109/tcds.2026.3735052.png)
Abstract
En 中文
As Transformer-based language models increasingly serve as repositories of factual knowledge, the ability to selectively forget specific information becomes crucial for correcting outdated facts, mitigating harmful content, or adhering to privacy regulations. However, current model editing methods either lack formal guarantees or incur undesirable interference with unrelated knowledge. In this work, we propose Provable Knowledge Forgetting (PKF), a rigorous framework that enables targeted suppression of factual information while preserving the model’s overall functionality. PKF formulates forgetting as a constrained optimization problem bounded by a Semantic Forgetting Bound, which quantifies the safe deviation from the model’s original behavior using Fisher Information. To make this tractable, we introduce a low-rank Fisher projection mechanism that ensures localized edits and a regularization term that minimizes semantic drift. We validate PKF on three benchmark datasets CounterFact, zsRE, and TriviaQA using five evaluation metrics: Forget@1, AvgKL, Knowledge Interference Score (KIS), KISnorm, and Neighborhood Consistency Rate (NCR). Our experiments show that PKF outperforms strong baselines such as ROME, MEND, and KE, achieving over 94% forgetting success with significantly lower interference. This work offers a scalable, theoretically grounded solution for controllable forgetting in language models, with implications for safety, compliance, and model lifespan management.
Keywords:
Knowledge forgetting
Transformer models
Model editing
Fisher information
Semantic drift
Journal
IF:
4.9
Papers:
1.0K
Citations:
3.5K
Organization
Cited Papers
No cited papers available

