Item

Multilingual Document Understanding: From Global Benchmarking to Arabic-Centric Evaluation

Heakl, Ahmed Mohamed Sobhy
Citations
Altmetric:
Supervisor
Department
Computer Vision
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
How well do modern document understanding systems truly comprehend the world’s written languages? Despite remarkable progress in vision-language models (VLMs), the ability to parse, recognize, and extract structured content from documents remains heavily skewed toward English and a handful of high-resource scripts. This thesis addresses the multilingual document understanding gap through two complementary contributions: a large-scale cross-lingual framework that exposes systemic failures across diverse writing systems, and a focused benchmark that dissects the unique challenges of Arabic, one of the most widely spoken yet computationally underserved languages. We first introduce DocAtlas, a framework for constructing high-fidelity OCR datasets and benchmarks covering 82 languages and 9 evaluation tasks through model-free annotation pipelines, yielding a 360K-page training corpus and a difficulty-stratified benchmark of 5,862 pages. Evaluating 16 state-of-the-art models reveals that low-resource scripts suffer 40–60% accuracy drops compared to high-resource counterparts, and that structured extraction plateaus at 71-73% TEDS regardless of language. We further demonstrate that Direct Preference Optimization (DPO) with rendering-derived ground truth achieves stable cross-lingual transfer, improving both in-domain (+1.9%) and out-of-domain (+1.8%) accuracy, where supervised fine-tuning degrades out-of-domain performance by up to 21%. Motivated by the persistent Arabic underperformance exposed by DocAtlas, we present KITAB-Bench, a comprehensive Arabic OCR benchmark spanning 9 domains and 36 sub-domains with 8,809 samples, covering tasks from basic text recognition to table extraction, chart understanding, diagram parsing, and end-to-end PDF-to-Markdown conversion. We introduce three evaluation metrics tailored to Arabic documents: MARS, CharTeX, and CODM. Our evaluation shows that modern VLMs outperform classical OCR by an average of 60% in character error rate, yet the best model achieves only 65% on PDF-to-Markdown conversion, exposing critical limitations in Arabic document understanding. Together, DocAtlas and KITAB-Bench provide a telescope-to-microscope view of multilingual document understanding: the former maps the global landscape, the latter dissects one of its most challenging scripts.
Citation
Heakl, Ahmed Mohamed Sobhy, "Multilingual Document Understanding: From Global Benchmarking to Arabic-Centric Evaluation," M.S. Thesis, Computer Vision, MBZUAI, 2026.
Source
Conference
Keywords
Arabic Document Understanding, Multilingual Document Understanding, Optical Character Recognition (OCR), Vision-Language Models, Direct Preference Optimization (DPO), Benchmark Evaluation
Subjects
Source
Publisher
DOI
Full-text link