Item

Exploring Practical LLM Capability Under Real-World Constraints: Data-Efficient Reasoning and Mobile-Centric Benchmarking

Bsharat, Sondos Mahmoud Yousef
Citations
Altmetric:
Department
Machine Learning
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Large language models (LLMs) have advanced rapidly, yet practical deployment still faces key challenges under real-world constraints. Reliable multi-step reasoning often requires costly supervision, and standard evaluation protocols are not well aligned with mobile settings, where queries and system constraints differ substantially from desktop or cloud environments. This thesis studies practical LLM capability through two complementary contributions: data-efficient reasoning improvement and mobile-centric evaluation. First, we introduce PSS (Principle-Guided Prompt-Space Scaling), a data-efficient method for improving reasoning by scaling the prompt space instead of relying on large curated datasets. Starting from 90 carefully selected seed problems, PSS applies principle-guided instructional wrappers to elicit diverse teacher rationales, producing compact training sets of up to 900 examples. Fine-tuning Qwen-2.5 models on this supervision yields consistent gains on reasoning benchmarks (AIME25, MATH500, GPQA-Diamond) and improves out-of-domain generalization. Second, we propose the Mobile-MMLU benchmark family to evaluate LLM capabilities in mobile-centric scenarios. Mobile-MMLU contains 16,186 multiple-choice questions across 80 practical fields and is designed for artifact-aware, order-invariant evaluation. We further construct Mobile-MMLU-Pro, a more challenging subset of 9,497 questions constructed via multi-model consistency and difficulty filtering, and introduce an open-style question (OSQ) setting that removes multiple-choice cues. Experiments show that mobile-oriented evaluation can change model rankings relative to traditional benchmarks and more clearly separate performance among smaller models relevant to on-device deployment. Together, these contributions improve and evaluate LLM capability under realistic deployment constraints by enabling data-efficient reasoning and mobile-centric evaluation.
Citation
Bsharat, Sondos Mahmoud Yousef, "Exploring Practical LLM Capability Under Real-World Constraints: Data-Efficient Reasoning and Mobile-Centric Benchmarking," M.S. Thesis, Machine Learning, MBZUAI, 2026.
Source
Conference
Keywords
LLM capability, Large language models, Reasoning, Data-efficient, Mobile-centric, Evaluation
Subjects
Source
Publisher
DOI
Full-text link