Rethinking Deep Learning Architectures for Dense Prediction Tasks in Medical Imaging
Qin, Chao
Qin, Chao
Citations
Altmetric:
Author
Supervisor
Department
Computer Vision
Embargo End Date
Type
Dissertation
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Deep learning has achieved remarkable success in dense prediction tasks such as object detection and segmentation, driven by advances in architectures including vision transformers (ViTs) and promptable foundation models. These models effectively encode long-range contextual dependencies and fine-grained representations. However, translating such architectures to medical imaging presents distinct challenges. Medical data span heterogeneous modalities, including ultrasound, CT, MRI, and histopathology, generally characterized by limited annotated data, substantial domain shifts, extreme scale variations (from cellular structures to whole organs), efficiency requirements for real-time clinical workflows, and the need for interactive refinement in dense scenarios such as cell segmentation in computational pathology. This thesis studies designing deep learning architectures for dense prediction in medical imaging by systematically addressing robustness, efficiency, and adaptability for medical modalities. First, a spatial-temporal deformable attention framework is introduced for breast lesion detection in ultrasound videos. By enabling deeper cross-frame feature aggregation while maintaining multi-frame inference, the proposed method mitigates instability caused by noise and motion artifacts, thereby improving lesion detection robustness in dynamic ultrasound sequences. Second, an efficient real-time spatial-temporal transformer lesion detector is proposed that reduces redundant encoder computation and performs high-level temporal fusion to accurately detect breast lesions in ultrasound videos with desired latency. Third, a foundation model is developed that robustly captures multi-scale anatomical features across 2D images, videos, and 3D volumetric sequences for segmenting anything in ultrasound imaging, achieving promising zero-shot generalization across benchmarks. Fourth, a dual-branch encoder adapted segment anything model framework is introduced that effectively encodes high-and low-level domain-specific features for universal medical image segmentation, achieving impressive segmentation performance on 30 public medical datasets covering different 2D and 3D medical modalities. Finally, this thesis presents a unified segmentation framework for computational pathology that supports both automatic nucleus instance segmentation and iterative interactive refinement through a shared prompt interface and a multi-scale decoder, addressing the frequent merge/split issues in dense cellularity.
Citation
Qin, Chao, "Rethinking Deep Learning Architectures for Dense Prediction Tasks in Medical Imaging," PhD Dissertation, Computer Vision, MBZUAI, 2026.
Source
Conference
Keywords
Foundation Model, Deep Learning, Transformer, Ultrasound, Segment Anything Model, DETR
