Loading...
Do LLMs model human linguistic variation? A case study in Hindi-English Verb code-mixing
Choudhary, Mukund ; Jindal, Madhur ; Aeron, Gaurja ; Choudhury, Monojit
Choudhary, Mukund
Jindal, Madhur
Aeron, Gaurja
Choudhury, Monojit
Files
Supervisor
Department
Natural Language Processing
Embargo End Date
Type
Conference proceeding
Date
License
http://creativecommons.org/licenses/by/4.0/
Language
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Do large language models (LLMs) model linguistic variation? We investigate this question through Hindi-English (Hinglish) verb code-mixing, where speakers can use either a Hindi verb or an English verb with the light verb karna (’do’). Both forms are grammatical, but speakers show unexplained variation in language choice for the verb. We compare human preferences on controlled code-mixed minimal pairs to LLM perplexities spanning families, sizes, and training language compositions. We find that current LLMs do not reliably classify verb language preferences to match native speaker judgments. We also see that with specific supervision, some models do predict human preference to an extent. We release native speaker acceptability judgments on 30 verb pairs, perplexity ratios for 4,279 verb pairs across 7 models, and experimental materials.
Citation
M. Choudhary, M. Jindal, G. Aeron, M. Choudhury, "Do LLMs model human linguistic variation? A case study in Hindi-English Verb code-mixing," 2026, pp. 5491-5509.
Source
Findings of the Association for Computational Linguistics: EACL 2026
Conference
Findings of the Association for Computational Linguistics: EACL 2026
Keywords
Subjects
Source
Findings of the Association for Computational Linguistics: EACL 2026
Publisher
Association for Computational Linguistics
