Item

Stability Regimes in Multi-Agent Large Language Model Coordination: A Cross-Scale Empirical Benchmark

Aldahmani, Mohammed Abdulla Mohammed
Citations
Altmetric:
Department
Machine Learning
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Multi-agent large language model (LLM) systems are often motivated by the intuition that coordinating multiple agents through voting, critique, or verification should improve reasoning quality. In practice, however, these coordination patterns can either improve accuracy or introduce systematic failure modes, and there is limited empirical guidance on when each design is reliable. This thesis presents a cross-task empirical study of multiagent coordination and interprets the resulting patterns as regime-like behavior rather than as evidence for a universal coordination law. The study compares four baseline coordination architectures—Single Agent, Ensemble Voting, Sequential Analyzer–Reasoner–Verifier, and Debate—together with two targeted extensions, MAD Symmetric and CoVe Sequential. Experiments cover five completed scales (8B, 14B, 20B, 32B, and 120B) on GSM8K, MMLU, ARC, and FEVER. In addition to accuracy and token cost, the analysis uses trace-level answer-change diagnostics to measure whether coordination-induced changes are more likely to fix errors than break correct answers. Results indicate that vote-based aggregation is the most robust accuracy-seeking default in the completed evidence base. Ensemble Voting is positive or near-neutral in almost all completed settings and avoids the large failure modes observed in critique- and verification-heavy protocols. Baseline Debate is frequently unstable, especially on arithmetic and abstention-sensitive verification. MAD Symmetric can outperform simpler baselines in selected knowledge-heavy settings but remains task-conditional and expensive. Verification is also conditional: Sequential can help in selected settings, whereas CoVe is often less reliable and sometimes destructive. To test whether the same coordination story changes under stronger external feed-back, the thesis also includes a dedicated HumanEval study under a one-repair open-test protocol. In that setting, executable feedback improves performance across five tested backbones, Single + feedback becomes a strong baseline, and the preferred coordination mechanism under feedback varies by model backbone: Debate is strongest on gpt-oss-20b and qwen3-32b, whereas Sequential is strongest on the other three tested backbones. A final post-hoc decision layer treats coordination as selective inference-time compute characterized by observed stability, predictability, and cost. Under that view, Ensemble is usually a robust default, Debate, Sequential, and MAD are better treated as selective mechanisms, and CoVe is harder to justify. The thesis contributes a cross-scale, cross-task empirical benchmark of major multi- agent coordination mechanisms, a trace-level diagnostic that distinguishes helpful from harmful coordination, a dedicated execution-feedback study on HumanEval, a cautious regime-like interpretation of the observed patterns, and an empirical decision framework that treats coordination as selective inference-time compute.
Citation
Aldahmani, Mohammed Abdulla Mohammed, " Stability Regimes in Multi-Agent Large Language Model Coordination: A Cross-Scale Empirical Benchmark," M.S. Thesis, Machine Learning, MBZUAI, 2026.
Source
Conference
Keywords
LLM Benchmarking, Multi-Agent LLMs, Coordination Stability
Subjects
Source
Publisher
DOI
Full-text link