Reinforcement Learning from Imperfect Data: Policy Extraction, Optimization, and Applications
Gao, Chengqian
Gao, Chengqian
Citations
Altmetric:
Author
Supervisor
Department
Machine Learning
Embargo End Date
Type
Dissertation
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Reinforcement learning (RL) holds the promise of iteratively improving policies through interaction with an environment or through learning from historical data. However, realizing this promise in practice remains challenging due to gaps between the idealized assumptions underlying theoretical guarantees and the constraints of real-world data collection. This thesis addresses four forms of data imperfection that arise across key components of the RL pipeline and develops principled algorithmic responses to each.
• Non-expert behavioural policies. Offline RL with mixed demonstrations suffers from catastrophic failure when policy improvement exploits out-of-distribution Q-predictions, a risk we identify and formalize. Enforcing Lipschitz continuity on the learned Q-function stabilizes value extrapolation, enabling existing algorithms such as TD3+BC and BEAR to recover expert-level performance even when 70% of training data come from random behavioural policies.
• Task-irrelevant observations. Redundant input features inflate model size and degrade generalization, a problem that we show is particularly severe for natural evolution strategies under high-dimensional inputs. Introducing a hard-thresholding operator with formal convergence guarantees, we develop NESHT, which achieves state-of-the-art performance even when 80% of observations are task-irrelevant Gaussian noise.
• Capacity-mismatched training data. Large language model (LLM) alignment commonly assumes all preference data contribute positively, regardless of model capacity. We identify a previously unexamined risk: preference data vary in difficulty, and overly difficult examples hinder alignment by exceeding the model’s representational capacity. Our method Selective DPO filters out such examples, achieving up to 16% higher win rates over standard DPO across multiple models and benchmarks.
• Conflicting optimization objectives. Concise reasoning in LLMs requires balancing accuracy, driven by RL, and brevity, enforced through heuristic reward shaping, a combination that lacks a principled foundation and yields unstable trade-offs. Reframing the problem as a Lagrangian min–max optimization, we propose PALU, which reduces response length by 64% while improving accuracy by 16% across six benchmarks, with robust generalization across model scales and reasoning domains.
Together, these contributions demonstrate that the gap between the promise of RL and its practical performance often arises from unmodeled data imperfections. By systematically characterizing and addressing these imperfections, this thesis shows that principled algorithmic treatment can yield robust policy improvement across diverse real-world settings, from locomotion control and visual game playing to LLM alignment and reasoning.
Citation
Gao, Chengqian, "Reinforcement Learning from Imperfect Data: Policy Extraction, Optimization, and Applications," PhD Dissertation, Machine Learning, MBZUAI, 2026.
Source
Conference
Keywords
offline reinforcement learning, reinforcement learning, large language model, data imperfection
