Return
LLM-generated scientific content detection method
DOI:10.1007/s00521-026-12465-6.png)
Abstract
En 中文
Large language models (LLMs), such as ChatGPT, Gemini, and LLaMA-3, can generate scientific text that closely resembles human-written text. While they simplify the creation of research papers and articles, they also raise concerns about plagiarism, misinformation, and academic dishonesty. Detecting AI-generated content is therefore crucial for maintaining scientific integrity. This study introduces the Artificial Intelligence Scientific Text Detector (AISciDetector), a hybrid model combining BigBird embeddings with a CNN-BiLSTM framework. BigBird captures long-range dependencies in scientific texts. The CNN extracts local features, and the BiLSTM models sequential patterns in both directions. The model performs binary classification to distinguish human-written from LLM-generated text. We also present LLMSciTxt, a dataset containing scientific texts by human authors and outputs from ChatGPT, Gemini, and LLaMA-3. Experiments show that AISciDetector achieves an accuracy of 91.5% and an F1-score of 91.49%, outperforming classical machine learning and transformer-based models. These results demonstrate that AISciDetector can effectively identify LLM-generated content, supporting the integrity and credibility of academic publications.
Keywords:
Large language models
Scientific content
Academic plagiarism detection
Transformer
Deep learning
Journal
IF:
4.5
Papers:
863
Citations:
3.2W
Organization
Cited Papers
Robust Natural Language Processing: Recent Advances, Challenges, and Future Directions
IEEE ACCESS
IF3.6

