Protocol-Controlled Zero-Shot Benchmarking of Vision–Language Models for Surgical VQA
Almahri, Muhra Salem Mohamed
Almahri, Muhra Salem Mohamed
Citations
Altmetric:
Author
Supervisor
Department
Machine Learning
Embargo End Date
2027-05-01
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Vision–language models (VLMs) are increasingly proposed for surgical visual question answering (VQA), yet reported accuracies vary widely across studies, making it difficult to assess what these models can actually do. Performance is highly sensitive to evaluation design, including answer format, contextual prompting, and the use of external knowledge. In this thesis, we present a protocol-controlled study of zero-shot surgical VQA to systematically examine how evaluation choices affect perceived model performance. We consider three commonly used evaluation protocols open-ended answering, contextual prompting, and multiple-choice (MCQ) selection and evaluate VLMs of different scales on two complementary datasets, EndoVis-18-VQA (robotic/laparoscopic surgery) and Kvasir-VQA (GI endoscopy). Beyond protocol effects, we investigate a range of knowledge integration strategies, including reasoning-oriented prompting, retrieval-augmented generation, and multimodal context construction. We analyze their impact across model scales and task types, highlighting when additional knowledge improves performance and when it introduces conflicts with visual evidence.Our results reveal that evaluation protocols can substantially alter performance rankings, with MCQ settings often inflating accuracy due to constrained answer spaces and option biases. We further identify a visual–knowledge trade-off: while knowledge augmentation benefits knowledge-intensive questions, it can degrade performance on visually grounded tasks when textual context is misleading. Finally, we conduct a reliability-aware analysis of failure modes, showing that knowledge-augmented strategies can lead to detrimental answer flips and overconfident errors, and we evaluate mitigation approaches such as selective retrieval and reranking.Our findings argue for protocol-controlled benchmarks and selective knowledge integration as prerequisites for trustworthy multimodal systems in surgical settings.
Citation
Almahri, Muhra Salem Mohamed, "Protocol-Controlled Zero-Shot Benchmarking of Vision–Language Models for Surgical VQA," M.S. Thesis, Machine Learning, MBZUAI, 2026.
Source
Conference
Keywords
retrieval-augmented generation, surgical VQA, vision–language models, medical AI safety, multimodal grounding
