From Pleural Line Segmentation to Video-Level Diagnosis: Deep Learning for Lung Ultrasound
Almsouti, Alya
Almsouti, Alya
Citations
Altmetric:
Author
Supervisor
Department
Computer Vision
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Patients with acute kidney injury (AKI), chronic kidney disease (CKD), and dialysis-dependent renal failure are at high risk of fluid overload and pulmonary edema. Early detection of extravascular lung water is clinically important because pulmonary congestion may progress before obvious respiratory symptoms appear. Lung ultrasound (LUS) is a practical bedside modality for this purpose, but interpretation remains operator-dependent and challenging, particularly for novice sonographers and in variable acquisition conditions. These challenges motivate the need for automated, reliable computer-aided analysis to support consistent assessment.
This thesis presents a deep learning framework for multi-class LUS video classification across four clinically relevant categories: healthy, B-lines, consolidation, and combined B-lines with consolidation. Using an open-access dataset of 1,886 videos from 219 patients, we implemented a patient-level preprocessing and five-fold cross-validation pipeline to avoid data leakage and improve class balance. A systematic baseline benchmark across encoder--aggregator combinations showed that an ImageNet-pretrained ViT-Small frame encoder with transformer-based temporal aggregation achieved the strongest baseline performance, with macro F1 of 0.628 ± 0.032.
Building on this baseline, this thesis investigates several methodological extensions, with a particular emphasis on anatomy-aware learning. Specifically, we evaluated: (i) Radon-feature integration for global line-structure cues, (ii) hard and soft hierarchical classification strategies, (iii) multi-task and pretraining schemes using pleural-line segmentation, and (iv) mask-guided attention supervision to encourage anatomically meaningful focus. Although Radon transforms are effective classical image-processing descriptors, they did not provide meaningful additional discrimination in this setting. Hierarchical formulations improved performance over flat classification, with one-versus-all reaching 0.645 ± 0.016 and soft hierarchy reaching 0.634 ± 0.030 macro F1. Segmentation-informed learning further improved a ResNet-based classifier from 0.560 to 0.609 macro F1, while pleural-line segmentation achieved Dice up to 0.816 under guided prompting. Most importantly, mask-guided attention supervision yielded the strongest overall classification result, reaching 0.657 ± 0.024 macro F1 (approximately +3\% absolute over the baseline) while improving attention localization around pleural-line regions.
Overall, the results indicate that structure-aware supervision and anatomy-guided representation learning improve both pathological separability and interpretability in LUS
video analysis. Beyond predictive performance, this work contributes a reproducible evaluation pipeline and a practical annotation workflow video segmentation tasks.
Citation
Almsouti, Alya, "From Pleural Line Segmentation to Video-Level Diagnosis: Deep Learning for Lung Ultrasound," M.S. Thesis, Computer Vision, MBZUAI, 2026.
Source
Conference
Keywords
Hierarchical learning, Lung ultrasound, Multi-class classification, Attention supervision, Pulmonary edema
