Towards Automated Evaluation of Citation Usage
Al Sikaiti, Iman Andrea King
Al Sikaiti, Iman Andrea King
Citations
Altmetric:
Author
Supervisor
Department
Natural Language Processing
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
The Related Work section plays a central role in academic writing of positioning a study within existing literature and demonstrating the author’s understanding of prior work. Despite this, most research has focused on automatically generating Related Work sections or classifying citation intent, with little attention paid to detecting writing issues within them. This thesis addresses that gap by treating the problem as span-level error detection in Introduction and Related Work sections. We propose an error taxonomy of four categories: Format, Unsupported Claim, Coherence, and Lacks Synthesis, grounded in analyses of the Exposia and NLPeer datasets and prior work on citation analysis and scientific discourse. Using this taxonomy, we construct a gold-standard dataset of 88 annotated documents from NLPeer, containing 188 manually
labelled error spans. To support model training, we also create a silver-standard dataset of 1,100 synthetic spans generated from unused NLPeer documents using GPT-5, with generation guided by span length and frequency distributions observed in the annotated data. We evaluated both prompt-based and encoder-based approaches on this dataset. The prompting baseline uses zero-shot and few-shot GPT-5 prompts, while encoder models include SciBERT, XLM-RoBERTa, ModernBERT, and NeoBERT. For short Unsupported Claim spans, a citation-based post-processing pipeline combining scientific entity recognition and citation filtering is also introduced. Models are evaluated using token-level F1, Cohen’s κ, and Pairwise F1, while synthetic data quality is assessed through structural and lexical analyses covering span length, positional distributions, MATTR, EAD-2, Top-20 token coverage, and vocabulary per span. Encoder models substantially outperform prompting baselines, particularly in localising span boundaries. Surface-level categories such as Format and short Unsupported Claim are detected most reliably, while discourse-level categories including Lacks Synthesis, and long Unsupported Claim remain difficult. Synthetic spans closely match real spans structurally but differ lexically, showing reduced diversity and repetitive phrasing. Correlation analysis further shows that lexical concentration in synthetic data is strongly linked to lower model performances, highlighting synthetic data quality as an important factor in performance. Overall, this thesis shows that error detection in Related Work sections can be effectively framed as a span-level task, but that performance depends on the quality of synthetic training data and the complexity of the targeted error type. This work contributes a new annotated dataset and synthetic data, an error taxonomy, and empirical findings on the relationship between citation reasoning, discourse structure, and automated writing quality assessment.
Citation
Al Sikaiti, Iman Andrea King, "Towards Automated Evaluation of Citation Usage," M.S. Thesis, Natural Language Processing, MBZUAI, 2026.
Source
Conference
Keywords
Error taxonomy, Citation Usage, Related Work, Argumentation, AI for Science, Span-level detection
