Item

Towards Robust Detection of AI-Generated Code

Orel, Daniil
Citations
Altmetric:
Department
Natural Language Processing
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Large Language Models (LLMs) have transformed code generation, offering unprecedented efficiency, but raising critical concerns for programming education, ethics, security, and assessment integrity. Detecting AI-generated code has thus become an urgent task. Prior research, however, has been limited in scope, often focusing on narrow domains, a small set of programming languages, and non-adversarial settings, leaving detectors fragile in real-world use. In this work, we present DroidCollection, the most comprehensive open dataset for machine-generated code detection, comprising over one million code samples across seven programming languages, 43 generation models, and multiple real-world domains. Beyond fully synthetic code, DroidCollection includes human–AI co-authored code and adversarially crafted samples designed to evade detection. Leveraging this dataset, we introduce DroidDetect, a suite of encoder-based detectors trained with a multi-task learning objective. Our framework is systematically evaluated across multi-lingual, multi-model, multi-domain, hybrid, and adversarial scenarios, including generalization to unseen languages, models, and domains. We benchmark classical machine learning models, pre-trained language models (PLMs), and LLM-based detectors, and further investigate techniques such as metric learning and uncertainty-based resampling to enhance robustness against noisy and adversarial distributions. Extensive experiments demonstrate that while existing detectors fail to generalize, DroidDetect consistently outperforms baselines, achieving strong robustness against adversarial attacks and hybrid authorship. This work establishes a new benchmark for AI-generated code detection, combining a robust model with a large-scale dataset, and provides the research community with the largest open foundation for studying authorship verification in programming.
Citation
Orel, Daniil, "Towards Robust Detection of AI-Generated Code," M.S. Thesis, Natural Language Processing, MBZUAI, 2026.
Source
Conference
Keywords
Code-LMs, LLMs, AI-generate content detection
Subjects
Source
Publisher
DOI
Full-text link