arrow
Return

Highly Generalizable Cross-Domain Machine-Generated Text Detection

delete2026-03-25
delete0
PRE
AI
Y
Yikang Xing
Y
Yuanhao Men
S
Senlin Luo
Z
Zongyuan Yang
J
Jinjie Zhou
潘丽敏 (Limin Pan) *
DOI:10.1016/j.neucom.2026.133426delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The machine-generated text has significant impacts on the reliability of information, leading to a series of technical and ethical issues. Existing machine-generated text detection methods are typically limited to specific domains and exhibit low accuracy when applied to cross-domain scenarios. Moreover, machine-generated texts with a low revision ratio often exhibit high semantic similarity to human-written texts, making them prone to misclassification by detection models. This study proposes a highly Generalizable cross-domain machine-generated text detection method (HGCD). A framework is constructed that synergistically leverages domain-specific and domain-general encoders, reinforcing the general features of machine-generated texts, thereby promoting cross-domain detection accuracy. In addition, an edit distance loss is designed to mitigate feature overlap between machine-generated and human-written texts, thus reducing the false detection rate. Experimental results show that, once trained on a single domain, HGCD achieves high detection accuracy in other domains without any fine-tuning, indicating its strong generalizability and broad applicability. Compared with the SOTA method, HGCD improves accuracy by up to 1.54% in in-domain scenarios and by up to 6.59% in cross-domain scenarios.
Keywords:
Machine-generated text detection
Cross-domain generalization
Edit distance loss
Domain-specific encoder
Domain-general encoder

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

No organization information available