Item

CopyShield: Comparing Copyright Defense Mechanisms for Large Language Models

Alshehyari, Maryam Salem Obaid
Citations
Altmetric:
Supervisor
Department
Machine Learning
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Large language models (LLMs) trained on massive corpora have demonstrated a well-documented tendency to memorize and reproduce portions of their training data verbatim, including copyrighted literary works. This phenomenon poses significant legal and ethical risks, as verbatim or near-verbatim reproduction of protected text may constitute copyright infringement under intellectual property law. While several defense mechanisms have been proposed in the literature, they have been developed and evaluated in isolation, making direct comparison difficult. This thesis presents CopyShield, a systematic study of three defense approaches that intervene at fundamentally different levels of the LLM generation pipeline to suppress copyright leakage while preserving model utility. Using a controlled experimental setup in which a LLaMA-3.1-8B base model is fine-tuned to memorize five public domain books as a proxy for copyrighted content, we evaluate each method under identical conditions against two co-equal objectives: copyright compliance and utility preservation. The three methods examined are: (1) contrastive decoding, which operates at the output level by down-weighting token probabilities that are elevated relative to a reference model trained without the target books; (2) Direct Preference Optimization (DPO), which operates at the behavioral level by fine-tuning the model to prefer non-reproducing responses; and (3) activation intervention, which operates at the representation level by training a linear classifier on hidden-state activations to detect and intercept copyright-reproducing generation before output is produced. Our evaluation employs a multi-metric framework covering both literal leakage (NV-Recall, Phase-1 LCS, ROUGE-L) and non-literal leakage (calibrated embedding similarity), alongside utility metrics including AI-judge scores for helpfulness, coherence, informativeness, and correctness, as well as refusal rate and degeneracy rate. Results show that all three methods reduce copyright leakage relative to the undefended baseline, but with distinct trade-off profiles. Contrastive decoding reduces mean NV-Recall by 23--27% and modestly reduces non-literal flagging across all lambda values (3.5--4.0% vs. 5.0% baseline), but preserves only 78.8--82.1% of raw base QA capability. DPO achieves the strongest literal suppression---eliminating all perfect copies and reducing mean NV-Recall by 99.2% to base model level (0.002)---while achieving the highest QA utility of all methods (99.4% of raw base capability) and introducing zero refusals; non-literal flagging remains comparable to the baseline (5.5%). Activation intervention reduces mean NV-Recall by 89% and achieves the lowest non-literal flagging (1/200, 0.5%) via an 84% intervention rate, but introduces a 6% false-positive rate on factual queries and preserves 94.9% of raw base QA capability. No single method dominates all dimensions, and the optimal choice depends on deployment context. The main contributions of this thesis are: (i) a comparative study of three defense approaches that span output, behavioral, and representation levels; (ii) a systematic and reproducible evaluation under identical conditions using statistically calibrated thresholds; and (iii) empirical evidence that the level of intervention determines the type and magnitude of the compliance--utility trade-off, providing actionable guidance for practitioners deploying LLMs under copyright constraints.
Citation
Alshehyari, Maryam Salem Obaid, "CopyShield: Comparing Copyright Defense Mechanisms for Large Language Models," M.S. Thesis, Machine Learning, MBZUAI, 2026.
Source
Conference
Keywords
Memorization, Large Language Models (LLMs), Copyright Leakage, Contrastive Decoding, Direct Preference Optimization (DPO), Activation Intervention
Subjects
Source
Publisher
DOI
Full-text link