Abnormality-Centric Echocardiography Decision Support: Temporal Modeling, Representation Learning, and Reasoning-Enabled Reporting
Taratynova, Darya
Taratynova, Darya
Citations
Altmetric:
Author
Supervisor
Department
Machine Learning
Embargo End Date
2027-05-01
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Echocardiography is the primary non-invasive modality for cardiac assessment, yet reliable automated analysis remains challenging due to the dynamic nature of the heart, ultrasound artefacts, and the complexity of clinical interpretation across diverse populations and diagnostic objectives. This thesis addresses four interconnected problems: multi-class CHD classification from fetal echocardiographic video, systematic evaluation of echocardiography foundation models, view-consistent video-text representation learning, and multi-view reasoning with structured report generation.
Temporal Prompt Alignment (TPA) introduces a framework for fetal ultrasound video classification combining vision-language foundation models with prompt-guided contrastive learning. By aligning temporal video embeddings with class-specific clinical text prompts, TPA captures motion-dependent cardiac dynamics and supports multi-class CHD recognition. A Conditional Variational Autoencoder with Style Modulation (CVAE-SM) provides calibrated video-level confidence estimates, consistently outperforming temporal aggregation baselines on a private fetal dataset.
CardioBench establishes the first comprehensive public benchmark for echocardiography foundation models, unifying eight datasets into a standardised evaluation suite spanning regression, classification, and view recognition. Systematic evaluation under zero-shot, linear probing, and alignment protocols reveals that temporal modeling is indispensable for functional tasks, domain-specific text encoders enforce physiologically meaningful cross-modal structure, and general-purpose encoders exhibit robustness under distribution shift yet fail to organise clinical signals for fine-grained reasoning.
EchoVision addresses the supervision mismatch in echocardiography pretraining, where study-level reports describe structures not observable in any single view. A view-structure visibility matrix with stochastic gating constructs view-consistent text supervision, reducing label noise while preserving clinically meaningful language. Trained with symmetric contrastive learning, EchoVision shows notable improvements on view-localised pathologies across public benchmarks.
EchoSonar-R introduces the first multi-view reasoning-enabled model for echocardiographic disease classification and report generation. A spatiotemporal video encoder combined with a structure-aware cardiac detector provides global motion dynamics and spatially grounded anatomical cues to a language model. Two-stage training, supervised fine-tuning followed by group relative policy optimisation, enables interpretable chain-of-thought reasoning while optimising for clinical correctness, improving macro balanced accuracy by 17.1% over the strongest baseline and halving mean clinically significant errors in generated reports.
Collectively, these contributions advance echocardiographic Artificial Intelligence from prenatal anomaly detection and foundation model benchmarking to view-consistent representation learning and reasoning-enabled report generation, demonstrating that reliable cardiac analysis requires temporal awareness, clinical language grounding, and interpretable multi-view reasoning.
Citation
Taratynova, Darya, "Abnormality-Centric Echocardiography Decision Support: Temporal Modeling, Representation Learning, and Reasoning-Enabled Reporting," M.S. Thesis, Machine Learning, MBZUAI, 2026.
Source
Conference
Keywords
Vision-Language Models, Echocardiography, Foundation Models, Medical Report Generation, Multimodal Representation Learning
