Detection of AI-Generated Text in Academic Writing Using Transformer-Based Models
DOI:
https://doi.org/10.64803/jocsaic.v3i1.178Kata Kunci:
Artificial Intelligence, Academic Writing, AI-Generated Text Detection, Transformer Models, Text ClassificationAbstrak
The rapid advancement of generative artificial intelligence (AI), particularly large language models (LLMs), has significantly transformed academic writing by enabling the automatic generation of coherent and contextually relevant text. While these technologies improve writing efficiency and accessibility, they also raise serious concerns regarding academic integrity, originality, and authorship. Conventional plagiarism detection tools are ineffective in identifying AI-generated content because such text is often original rather than copied. This study proposes a transformer-based approach for detecting AI-generated text in academic writing by evaluating the performance of four pre-trained language models: BERT, RoBERTa, DistilBERT, and DeBERTa. The research methodology consists of dataset collection, text preprocessing, dataset splitting, transformer model fine-tuning, and performance evaluation using Accuracy, Precision, Recall, F1-score, and ROC-AUC. Experimental results show that all evaluated transformer models achieved excellent classification performance, with DeBERTa producing the highest accuracy of 98.10%, followed by RoBERTa (97.20%), BERT (95.10%), and DistilBERT (93.40%). These findings demonstrate that transformer-based architectures effectively capture contextual and semantic characteristics that distinguish AI-generated text from human-authored academic writing. The proposed approach provides a reliable solution for supporting academic integrity and assisting educators, publishers, and research institutions in detecting AI-generated content within scholarly documents.
Referensi
[1] Vera E. Woloshyn, Sam Illingworth, and Snežana Obradović-Ratković, “Introduction to Special Issue Expanding Landscapes of Academic Writing in Academia,” Brock Educ. J., vol. 33, no. 1, pp. 3–9, 2024, doi: 10.26522/brocked.v33i1.1118.
[2] J. Hutson, “Rethinking Plagiarism in the Era of Generative AI,” J. Intell. Commun., vol. 4, no. 1, 2024, doi: 10.54963/jic.v4i1.220.
[3] S. M. A. Shahid, M. N. Ali, M. H. Sarkar, and M. H. Rahman, “Ensuring Authenticity in Scientific Communication Approaches to Detect and Deter Plagiarism,” TAJ J. Teach. Assoc., vol. 37, no. 1, pp. i–iii, 2024, doi: 10.70818/taj.v37i1.0157.
[4] J. A. Oravec, “Artificial Intelligence Implications for Academic Cheating: Expanding the Dimensions of Responsible Human-AI Collaboration with ChatGPT and Bard,” J. Interact. Learn. Res., vol. 34, no. 2, pp. 213–237, 2023, doi: 10.70725/304731gmmvhw.
[5] K. C. Fraser, H. Dawkins, and S. Kiritchenko, “Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods,” J. Artif. Intell. Res., vol. 82, pp. 2233–2278, 2025, doi: 10.1613/jair.1.16665.
[6] V. B. Gupta, N. Chitranshi, V. Palanivel, S. Sheriff, D. Basavarajappa, and V. Gupta, “Critical Perspectives on Ethical Challenges in Higher Education: Analysing Contemporary Practices and Future Considerations,” J. Educ. Learn., vol. 14, no. 6, p. 119, 2025, doi: 10.5539/jel.v14n6p119.
[7] J. Q. J. Liu et al., “The great detectives: humans versus AI detectors in catching large language model-generated medical writing,” Int. J. Educ. Integr., vol. 20, no. 1, p. 8, 2024, doi: 10.1007/s40979-024-00155-6.
[8] V. H. Kalmani, A. C. Adamuthe, A. P. Gondil, V. P. Patil, R. A. Kore, and V. M. Metkari, “AI vs. Human Writing: Developing a Novel Method for Text Authenticity Detection in Education,” Int. J. Mod. Educ. Comput. Sci., vol. 17, no. 3, pp. 45–58, 2025, doi: 10.5815/ijmecs.2025.03.04.
[9] M. S. Islam, L. Xiangdong, and J. Ahmed, “BERT: advancements in language understanding for different NLP tasks: challenges and future perspectives,” J. Electr. Syst. Inf. Technol., vol. 13, no. 1, p. 49, 2026, doi: 10.1186/s43067-026-00341-1.
[10] A. Şentürk, A. Albayrak, and S. Arpacı, “Multi-Class News Classification with BERT, DistilBERT, RoBERTa, and ELECTRA Natural Language Processing Models,” Düzce Üniversitesi Bilim ve Teknol. Derg., vol. 14, no. 1, pp. 117–129, 2026, doi: 10.29130/dubited.1737003.
[11] A. Onan, “Hierarchical graph-based text classification framework with contextual node embedding and BERT-based dynamic fusion,” J. King Saud Univ. - Comput. Inf. Sci., vol. 35, no. 7, p. 101610, 2023, doi: 10.1016/j.jksuci.2023.101610.
[12] A. Jashari, “Enhancing Coherence and Persuasiveness in Academic Discourse through Lexical Cohesion and Critical Thinking,” Open J. Soc. Sci., vol. 14, no. 01, pp. 471–483, 2026, doi: 10.4236/jss.2026.141028.
[13] J. Al Hosni, “Preserving Authorial Voice in Academic Texts in the Age of Generative AI: A Thematic Literature Review,” Arab World English J., vol. 16, no. 3, pp. 244–258, 2025, doi: 10.24093/awej/vol16no3.14.
[14] X. Liang, X. Zhou, H. Zou, Y. Lu, and J. Qu, “DeepDiveAI: Identifying AI-Related Documents in Large Scale Literature Dataset,” J. Soc. Comput., vol. 6, no. 2, pp. 158–169, 2025, doi: 10.23919/JSC.2025.0007.
[15] M. M. Danyal, S. S. Khan, M. Khan, M. B. Ghaffar, B. Khan, and M. Arshad, “Sentiment Analysis Based on Performance of Linear Support Vector Machine and Multinomial Naïve Bayes Using Movie Reviews with Baseline Techniques,” J. Big Data, vol. 5, pp. 1–18, 2023, doi: 10.32604/jbd.2023.041319.
[16] E. Ferrara, “GenAI against humanity: nefarious applications of generative artificial intelligence and large language models,” J. Comput. Soc. Sci., vol. 7, no. 1, pp. 549–569, 2024, doi: 10.1007/s42001-024-00250-1.
Unduhan
Diterbitkan
Terbitan
Bagian
Lisensi
Hak Cipta (c) 2026 Cindy Atika Rizki, Nabila Khairuniza, Mel Wulandini (Author)

Artikel ini berlisensiCreative Commons Attribution-ShareAlike 4.0 International License.







