Item

Weak-to-Strong Verification: Benchmarking the Ability of Small Language Models to Generate Robust Contracts for Code Agents

Alzaabi, Abdalla Khalid Mohamed
Citations
Altmetric:
Supervisor
Department
Machine Learning
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
The emergence of the Internet of Agents (IoA) introduces a paradigm in which autonomous agents with heterogeneous capabilities collaborate and delegate tasks. In such settings, resource-constrained agents often rely on more capable agents to solve complex problems, giving rise to a fundamental trust challenge: a weak agent must verify whether a returned solution is correct, complete, and aligned with its intent. Inadequate verification mechanisms create opportunities for stronger agents to exploit this asymmetry by producing superficially plausible or minimal-effort solutions that satisfy weak checks without being truly correct. In this work, we investigate the verification gap in weak-to-strong delegation by systematically evaluating the ability of weak language models to generate lightweight verification contracts. We introduce a controlled benchmarking framework in which a weak model produces unit-test–style contracts for challenging programming tasks, derived from HumanEval-style benchmarks, while stronger models are explicitly optimized to generate token-efficient, adversarial solutions designed to satisfy these contracts. Crucially, correctness is evaluated using ground-truth execution against hidden test cases, enabling an objective assessment of whether the weak model’s accept or reject decisions are valid. Our results reveal that weak verification contracts are systematically exploitable, with False Accept where an incorrect solution is accepted emerging as the dominant failure mode across all experimental configurations. Comparison against an unguided baseline confirms that structured contract-based verification provides meaningful improvement over unin- formed decision making, yet remains fundamentally insufficient for dependable delegation under adversarial conditions. Root cause analysis identifies failures arising from two inter acting subsystems: the contract generation process, which produces incomplete, incorrectly specified, or structurally malformed assertions, and the judgment process, which remains poorly grounded in contract content across all three weak models evaluated. These findings demonstrate that weak agents cannot reliably serve as standalone verifiers of stronger agents’ outputs and underscore the need for improved contract generation, more reliable judgment mechanisms, and hybrid verification approaches in IoA-based systems.
Citation
Alzaabi, Abdalla Khalid Mohamed, "Weak-to-Strong Verification: Benchmarking the Ability of Small Language Models to Generate Robust Contracts for Code Agents," M.S. Thesis, Machine Learning, MBZUAI, 2026.
Source
Conference
Keywords
Agent Collaboration Protocols., Internet of Agents., Networked Agents and Decentralized AI.
Subjects
Source
Publisher
DOI
Full-text link