Item

From Conventional Deep Neural Networks to Retrieval-Grounded Vision-Language Modeling for Medical Diagnosis

Kassem, Mai Ahmed Shaaban Mohamed
Citations
Altmetric:
Department
Machine Learning
Embargo End Date
Type
Dissertation
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Artificial intelligence (AI) has become integral to medical diagnosis by augmenting clinical decision making, driven by the increasing availability of electronic health records (EHRs), medical imaging, and large-scale clinical text corpora. Recent advances in deep learning and vision-language models (VLMs) have demonstrated state-of-the-art performance across diagnostic tasks. However, reliable clinical deployment remains limited by several challenges: sensitivity to hyperparameter configurations, insufficient domain adaptation, heterogeneity and incompleteness of multimodal data, hallucinated outputs in generative models, and limited scalability across tasks and imaging modalities. This thesis addresses these challenges by developing a global optimization strategy to reduce hyperparameter sensitivity, domain-adapted language models for clinical text, grounded multimodal frameworks capable of handling heterogeneous and incomplete data, retrieval-augmented mechanisms to mitigate hallucinations, and parameter-efficient architectures to improve scalability across tasks and modalities. The thesis introduces several methodological contributions: OptBA formulates hyperparameter tuning as a population-based global optimization problem by integrating the Bees Algorithm with deep neural networks for medical text classification. Compared to conventional tuning strategies, this approach improves convergence stability, reduces susceptibility to suboptimal local minima, and enhances experimental reproducibility. LLM-Based Symptom Recognition adapts multilingual biomedical transformers to Spanish clinical narratives via supervised fine-tuning. This framework strengthens domain-specific named entity recognition in low-resource clinical settings and improves robustness against linguistic variability and ambiguity. MedPromptX enables multimodal diagnostic reasoning by integrating chest X-ray images with structured EHR data using visual grounding and dynamically refined few-shot prompting. The framework mitigates the impact of incomplete structured input and reduces the need for extensive retraining, supporting a reliable diagnosis under real-world conditions. MOTOR and MM-RAG-MedVQA address hallucinations in medical visual question answering through grounded multimodal retrieval and re-ranking. These approaches improve retrieval precision, answer faithfulness, and clinical relevance, validated on benchmark datasets and supported by expert evaluation. ME-VLIP introduces a modular and parameter-efficient vision-language architecture that combines quantized low-rank adaptation with dynamic task routing. The framework generalizes across multiple imaging modalities and diagnostic tasks without full model retraining, advancing scalable and generalist medical VLM design. These contributions articulate a methodological progression from conventional deep neural networks to retrieval-grounded vision-language modeling, advancing interpretable and scalable AI systems for medical diagnosis.
Citation
Kassem, Mai Ahmed Shaaban Mohamed, "From Conventional Deep Neural Networks to Retrieval-Grounded Vision-Language Modeling for Medical Diagnosis," PhD Dissertation, Machine Learning, MBZUAI, 2026.
Source
Conference
Keywords
Retrieval-Augmented Generation, Medical Visual Question Answering, Vision-Langauge Models
Subjects
Source
Publisher
DOI
Full-text link