Item

Stereotype Bias in Bilingual Setting: A Culturally Grounded Evaluation in Kazakhstan

Laiyk, Nurkhan
Citations
Altmetric:
Supervisor
Department
Natural Language Processing
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Stereotype bias in language models has been widely examined in English, but it remains largely understudied in bilingual contexts where multiple linguistic and cultural systems interact. This gap is especially important in regions where language use reflects complex historical and sociopolitical influences. This thesis focuses on Kazakhstan, a bilingual society where Kazakh, a low-resource Turkic language, and Russian, a high-resource Slavic language, are both actively used and frequently mixed through code-switching in everyday communication. We introduce Aqbileq, a high-quality, human-verified dataset consisting of 5,634 stereotype-bearing statements in Kazakh, Russian, and code-switched forms, covering six culturally salient domains that reflect the social realities of Kazakhstan, including cultural and geographic identity, demographics, ideology and religion, language use, social and family roles, and economic status. The dataset was constructed through a multi-stage pipeline involving template design by native Kazakh speakers, contrastive slot filling to produce stereotypical and counter-stereotypical pairs, manual code-switching, and human validation, with strong inter-annotator agreement achieved across all annotator pairs. We evaluate a total of nine language models, including three encoder only models and six decoder only large language models, covering both general multilingual architectures and Kazakh specific systems. Our evaluation combines perplexity-based scoring with a pretraining simulation of KazRoBERTa across twenty checkpoints to examine when and how stereotype bias emerges during training, as well as a generation based analysis that measures sentiment polarity in short Kazakh stories produced from stereotype-bearing prompts. Our findings indicate that stereotype bias is most pronounced in code-switched inputs, that Kazakh specific models tend to exhibit higher bias than their multilingual counterparts, and that the composition of pretraining data plays a substantial role in shaping how bias develops over time.
Citation
Laiyk, Nurkhan, "Stereotype Bias in Bilingual Setting: A Culturally Grounded Evaluation in Kazakhstan," M.S. Thesis, Natural Language Processing, MBZUAI, 2026.
Source
Conference
Keywords
large language models, stereotype bias, code-switching
Subjects
Source
Publisher
DOI
Full-text link