Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
Huy, Ta Duc ; Huynh, Duy Anh ; Xie, Yutong ; Qi, Yuankai ; Chen, Qi ; Le Nguyen, Phi ; Tran, Sen Kim ; Phung, Son Lam ; van den Hengel, Anton ; Liao, Zhibin ... show 3 more
Huy, Ta Duc
Huynh, Duy Anh
Xie, Yutong
Qi, Yuankai
Chen, Qi
Le Nguyen, Phi
Tran, Sen Kim
Phung, Son Lam
van den Hengel, Anton
Liao, Zhibin
Supervisor
Department
Computer Vision
Embargo End Date
Type
Conference proceeding
Date
License
Language
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Visual grounding (VG) is the capability to localize specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability and cross-modal alignment of multimodal models by highlighting relevant pathological features corresponding to textual descriptions. Current vision-language models struggle to associate textual descriptions with disease regions due to inefficient attention mechanisms and a lack of fine-grained token representations. In this paper, we empirically demonstrate two key observations. First, current VLMs assign high norms to background tokens, diverting the model's attention from regions of disease. Second, the global tokens used for cross-modal learning are not representative of local disease tokens. This hampers identifying correlations between the text and disease tokens. To address this, we introduce simple, yet effective Disease-Aware Prompting (DAP) process, which uses the explainability map of a VLM to identify the appropriate image features. This simple strategy amplifies disease-relevant regions while suppressing background interference. Without any additional pixel-level annotations, DAP improves visual grounding accuracy by 20.74% compared to state-of-the-art methods across three major chest $X$-ray datasets.
Citation
T.D. Huy, D.A. Huynh, Y. Xie, Y. Qi, Q. Chen, P. Le Nguyen , et al., "Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding," 2026, pp. 24445-24455.
Source
2025 IEEE/CVF International Conference on Computer Vision (ICCV)
Conference
2025 IEEE/CVF International Conference on Computer Vision (ICCV)
Keywords
46 Information and Computing Sciences, 4611 Machine Learning
Subjects
Source
2025 IEEE/CVF International Conference on Computer Vision (ICCV)
Publisher
IEEE
