Return
mailcom: Pseudonymization tool for textual data
DOI:10.1016/j.softx.2026.102703.png)
Abstract
En 中文
The rapid growth of data and its use has heightened concerns about data privacy. To support research on data that contains sensitive attributes, we developed the mailcom package for pseudonymization. mailcom is built entirely on open-source libraries and is designed for configurability and extensibility. Several languages are supported by default, with the option to add additional languages. Since pseudonymized outputs still require human review to guarantee full anonymization, the package serves as a scalable pre-processing layer that reduces manual work while establishing a principled baseline of privacy protection, enabling reproducible research on digital text under current data-ethics and governance standards.
Keywords:
Pseudonymization
Sensitive data
Natural language processing
Named entity recognition
Digital humanities

