Towards Interactive, Interpretable Multimodal Understanding and Reasoning with Foundation Models
Bangalath, Hanoona
Bangalath, Hanoona
Citations
Altmetric:
Author
Supervisor
Department
Computer Vision
Embargo End Date
Type
Dissertation
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Multimodal understanding is now shaped by Large Multimodal Models (MLLMs) that jointly model visual content and language to support open-ended interaction through natural language, moving beyond task-specific predictors toward general-purpose multimodal interaction. This shift also brings forth a persistent reliability gap: despite coherent and well-formed responses, MLLMs often fall back on superficial correlations and language priors rather than sustained engagement with visual evidence. The problem becomes more prominent as visual content and user query grow more complex, where models can drift from the input and follow textual shortcuts instead of staying anchored to the visual stream. This thesis takes a unified perspective on these limitations and argues that shortcut behavior is not only architectural, but also a data issue: training supervision shapes what models learn to attend to and rely upon. The thesis therefore treats interactability and interpretability as central requirements for closing the gap, because structured interaction and explicit grounding make evidence use both easier to sustain and easier to diagnose. Following a consistent loop of identifying capability gaps, diagnosing why existing supervision allows them to persist, and redesigning training data and evaluation protocols to directly incentivize grounded and interpretable behavior, the thesis progresses from images to videos and from static prompting to interactive reasoning. It first introduces GLaMM, a grounded conversation model that unifies scene level understanding, region-level interaction, and grounded responses within a single framework, and uses a large-scale, semi-automated data engine to generate supervision that keeps outputs explicitly tied to the visual scene. It next introduces PALO, a fully open-source MLLM spanning ten major languages, showing how data-centric multilingual instruction tuning broadens multimodal interaction beyond English without language-specific variants. For videos, where naturally aligned supervision is scarce and noisy, the thesis introduces a structured, data-centric video data engine that converts raw web videos into better-aligned video-text pairs and uses this refined supervision to train the Perception Encoder for temporally grounded video representations. To make multi-step reasoning limitations measurable, it introduces VideoMathQA, a benchmark of instructional videos with expert, timestamped reasoning traces that expose failures in integrating temporally distributed evidence across modalities. Finally, it introduces Video-CoM, an interactive video reasoning framework that shifts models from thinking about a video to thinking with the video through repeated re-engagement with the evolving visual stream, using data-centric supervision that forces evidence gathering so intermediate steps remain grounded rather than drifting in language space. Together, these contributions develop a unified data-centric approach to multimodal understanding, showing that supervision designed for grounded interaction and interpretability encourages models to remain anchored to visual evidence, improving reliability and performance.
Citation
Bangalath, Hanoona, "Towards Interactive, Interpretable Multimodal Understanding and Reasoning with Foundation Models," PhD Dissertation, Computer Vision, MBZUAI, 2026.
Source
Conference
Keywords
Pixel-level Grounding, Interactivity, Interpretability, Interactive Video Reasoning
