TiCO: Evidence-Driven Visual Reasoning via Token-Level Semantics
Zhao, Haoze
Zhao, Haoze
Citations
Altmetric:
Author
Supervisor
Department
Computer Vision
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Recent advances in multimodal large language models have shifted toward the “thinking with images” paradigm to improve reasoning faithfulness, yet existing methods rely heavily on explicit coordinate-based grounding. We identify two fundamental limitations in this approach, i.e., a semantic gap, where sparse regressions fail to capture high-dimensional latent structures, and a resolution gap caused by the granularity mismatch between patch-level tokens and pixel-level coordinates. To address these issues, we propose TiCO: Evidence-Driven Visual Reasoning via Token-Level Semantics (Thinking with Images without Coordinates). TiCO reformulates visual reasoning as a semantic evidence discovery process in latent space, where evidence is injected as token-level regions rather than pixel coordinates. Specifically, TiCO employs a dedicated probe token to estimate similarity distributions and performs importance sampling over spatially connected components to generate diverse and semantically dominant evidence rollouts. Extensive experiments on multiple high-resolution and general visual reasoning benchmarks demonstrate that TiCO consistently outperforms state-of-the-art coordinate-based methods while achieving significantly better inference efficiency.
Citation
Zhao, Haoze, "TiCO: Evidence-Driven Visual Reasoning via Token-Level Semantics," M.S. Thesis, Computer Vision, MBZUAI, 2026.
Source
Conference
Keywords
Thinking with Images, Visual Language Model, Visual Reasoning
