Item

Macaron: A Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling

Elsetohy, Alaa Ahmed Ibrahim
Citations
Altmetric:
Department
Natural Language Processing
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Evaluating whether large language models (LLMs) genuinely understand the world from a non-English perspective requires benchmarks that are simultaneously culturally authentic, reasoning-controlled, and cross-lingually comparable. Existing multilingual benchmarks either inherit English-centric cultural assumptions through translation or sacrifice cross-cultural control by authoring questions independently for each language, making it impossible to attribute performance differences to language, cultural knowledge, or reasoning difficulty. This thesis introduces Macaron, a template-first benchmark that factorises reasoning type and cultural aspect across question languages. The central methodological contribution is a set of 100 language-agnostic templates, each tagged with an explicit reasoning type and cultural aspect, which native annotators from 20 cultural communities instantiate with locally grounded content in both English and their local language. This design separates by construction the three factors that multilingual benchmarks typically conflate: the reasoning skill a question requires, the cultural aspect it probes, and the language in which it is presented. It thereby enables principled cross-cultural and cross-lingual comparison while maintaining cultural authenticity. Each cultural scenario yields six aligned evaluation instances: multiple-choice and True/False in both English and the local language. The resulting dataset contains 11,862 evaluation instances spanning 20 languages and dialects, 10 writing scripts, and 7 reasoning types across 22 cultural aspects. We evaluate 21 multilingual LLMs in a zero-shot setting, spanning closed-source thinking models, closed-source standard models, and open-weight instruction-tuned models ranging from 3B to 235B parameters. Our evaluation reveals four main findings. First, thinking models substantially outperform standard closed and open-weight models (80.8% vs. 74.8% and 58.0% overall accuracy), with open-weight models at the 3–8B scale approaching random performance on local-language multiple-choice. Second, English–local performance gaps concentrate in lower-resource languages (e.g., Amharic, Yoruba, Zulu, and Arabic dialects) and widen sharply with reduced model capacity, while China is the only context where local-language accuracy consistently matches English, attributable to the strong presence of Qwen models. Third, reasoning type imposes a consistent difficulty ordering independent of cultural context: mathematical and counting reasoning is the hardest type for 20 out of 21 models, while causal and commonsense reasoning are consistently easiest, reflecting the double burden of culture-specific numeric retrieval combined with arithmetic computation. Fourth, the paired True/False accuracy metric, which counts a scenario correct only if both the True and False verification questions are answered correctly, drops by an average of 29.2 percentage points from per-question accuracy, revealing that most models exploit positive-response bias rather than genuinely verifying cultural facts. Macaron provides the NLP community with a reproducible, extensible diagnostic tool for measuring culturally grounded reasoning across the diversity of human language and culture. The dataset is publicly available at https://huggingface.co/datasets/AlaaAhmed2444/Macaron.
Citation
Elsetohy, Alaa Ahmed Ibrahim, "Macaron: A Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling," M.S. Thesis, Natural Language Processing, MBZUAI, 2026.
Source
Conference
Keywords
evaluation, multilingual benchmark, datasets for low resource languages, NLP datasets, Multilingual Reasoning, Cultural evaluation
Subjects
Source
Publisher
DOI
Full-text link