ShariaBench: A Cross-Lingual Benchmark for Islamic Finance Knowledge and Reasoning
Togmanov, Mukhammed
Togmanov, Mukhammed
Citations
Altmetric:
Author
Supervisor
Department
Machine Learning
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Evaluating large language models (LLMs) on domain-specific knowledge and structured reasoning remains a critical challenge, particularly for fields governed by complex legal and ethical frameworks. Islamic finance operates under Shariah principles—including the pro- hibition of interest (riba), risk-sharing, and structured wealth redistribution—that impose fundamentally distinct knowledge and reasoning requirements compared to conventional finance. Despite serving over a billion speakers across diverse linguistic communities, no multilingual benchmark exists that evaluates LLMs on Islamic finance knowledge across languages, let alone on formal rule-bound reasoning under Shariah jurisprudence. To address this gap, we introduce ShariaBench, the first comprehensive multilingual benchmark unifying Islamic finance knowledge evaluation and verifiable symbolic reasoning across six languages. ShariaBench covers six core domains—Islamic banking, takaful, social finance and zakat, investment screening, capital markets, and Shariah-compliant contracts—and comprises two complementary components. The first is a multiple-choice question (MCQ) component combining translated and natively authored questions, en- abling direct comparison of cross-lingual transfer and culturally grounded financial knowl- edge. The second is a symbolic reasoning component of expert-validated calculation tem- plates grounded in formal Shariah rules, generating scalable, contamination-free, and fully verifiable multi-step reasoning instances across six calculation categories.
We evaluate a diverse set of open-source and proprietary LLMs under zero- to three- shot prompting settings, and augment the symbolic evaluation with finance-specialised and math-specialised models to probe the distinct contributions of domain knowledge and arithmetic reasoning ability. Our results reveal four key findings. First, model rankings invert between translated and native MCQ evaluation, exposing a translation dependency in open-source models that inflates performance on translated benchmarks while mask- ing weaker generalisation to natively phrased content. Second, native questions expose hidden difficulty that translated evaluation systematically fails to capture, particularly for morphologically complex languages where longer question stems cause sharp accuracy degradation. Third, confidence calibration reverses direction between evaluation condi- tions: models are overconfident on translated questions yet underconfident on native ones, revealing that models anchor their certainty to surface-level linguistic familiarity rather than genuine domain knowledge. Fourth, symbolic reasoning exposes a sharp dissocia- tion between declarative knowledge and rule-bound inference: math-specialised models substantially outperform larger general-purpose LLMs on Shariah calculation tasks, while boolean predicate evaluation emerges as the dominant failure mode regardless of model family or scale. Overall, ShariaBench exposes persistent weaknesses in multilingual and culturally grounded financial reasoning, and demonstrates that translated benchmarks systemati- cally underestimate both the difficulty and the calibration failure of current LLMs on Islamic finance. ShariaBench provides a rigorous, multi-track foundation for developing trustworthy, interpretable, and verifiable AI systems for global Islamic finance applications.
Citation
Togmanov, Mukhammed, "ShariaBench: A Cross-Lingual Benchmark for Islamic Finance Knowledge and Reasoning," M.S. Thesis, Machine Learning, MBZUAI, 2026.
Source
Conference
Keywords
Symbolic data, MMLU, AAOIFI, LLM, NLP, GRPO
