arrow
Return

Selective Forgetting in Document Images Using Enhanced Ensembles

delete2026-01-01
delete0
PRE
AI
M
Muhammad Mashhood
M
Momina Moetesum *
F
Faisal Shafait
A
Adnan Ul-Hasan
DOI:10.1007/978-3-032-04617-8_5delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Governments worldwide have enforced laws and regulations to protect digital data and ensure privacy. Data owners thus have been granted the right to be forgotten. An integral part of ensuring compliance requires companies to take reasonable steps to delete user data not only from databases but also from any AI models trained on it. In addition, model owners may also need to remove data harming the model's utility, and doing so efficiently, via Selective Forgetting, is crucial. Although extensively studied in classification, its application to document object detection remains underexplored. This paper introduces a novel selective unlearning framework using enhanced Sharded, Isolated, Sliced, and Aggregated (SISA) training combined with an Enhanced Weighted Box Fusion (WBF) strategy. By partitioning datasets into isolated shards and slices, we enable localized unlearning through selective sub-model retraining while mitigating performance loss via ensemble aggregation techniques. Experiments on the ICDAR 2017 POD and Invoices datasets demonstrate that our approach achieves better aggregation performance compared to standard WBF and soft Non-Maximum Suppression (sNMS), striking a balance between unlearning efficiency and model accuracy. These findings provide insights into scalable, privacy-preserving document AI for real-world applications.
Keywords:
Machine Unlearning
Document Analysis
Object Detection
Privacy Protection

Journal

D
DOCUMENT ANALYSIS AND RECOGNITION-ICDAR 2025, PT II
IF:
0
Papers:
14
Citations:
0

Organization

N
national university of sciences & technology - pakistan
Scholars:
7.8K
Papers: 6.6K
Citations: 6