Return
Highly Generalizable Cross-Domain Machine-Generated Text Detection
DOI:10.1016/j.neucom.2026.133426.png)
Abstract
En 中文
The machine-generated text has significant impacts on the reliability of information, leading to a series of technical and ethical issues. Existing machine-generated text detection methods are typically limited to specific domains and exhibit low accuracy when applied to cross-domain scenarios. Moreover, machine-generated texts with a low revision ratio often exhibit high semantic similarity to human-written texts, making them prone to misclassification by detection models. This study proposes a highly Generalizable cross-domain machine-generated text detection method (HGCD). A framework is constructed that synergistically leverages domain-specific and domain-general encoders, reinforcing the general features of machine-generated texts, thereby promoting cross-domain detection accuracy. In addition, an edit distance loss is designed to mitigate feature overlap between machine-generated and human-written texts, thus reducing the false detection rate. Experimental results show that, once trained on a single domain, HGCD achieves high detection accuracy in other domains without any fine-tuning, indicating its strong generalizability and broad applicability. Compared with the SOTA method, HGCD improves accuracy by up to 1.54% in in-domain scenarios and by up to 6.59% in cross-domain scenarios.
Keywords:
Machine-generated text detection
Cross-domain generalization
Edit distance loss
Domain-specific encoder
Domain-general encoder
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W
Organization
No organization information available

