Yuheng Li
BME PhD Defense Presentation
Date: 2026-09-03
Time: 9:30 AM EST
Location / Meeting Link: https://emory.zoom.us/meetings/97275597738/invitations?signature=C33oepLefYRvMkcT63UIIn6GhZwesuMD7UDVckh5D10
Committee Members:
Xiaofeng Yang, Ph.D (advisor); Wang Yun, Ph.D; Judy Wawira Gichoya, MD; John Oshinski, Ph.D; Xiao Hu, Ph.D
Title: Learning transferable CT representations for radiology and radiation oncology
Abstract:
Computed tomography (CT) is central to radiology and radiation oncology. In diagnostic imaging, CT supports the detection, localization, and characterization of disease. In radiation oncology, it provides the anatomical basis for organ-at-risk (OAR) and tumor delineation, image registration, treatment planning, dose assessment, and longitudinal response evaluation. These applications require models that can represent both normal anatomy and pathological findings across anatomical regions, imaging protocols and contrast phases, and disease populations. However, many existing artificial intelligence methods for CT are developed and evaluated for individual tasks, a single acquisition or contrast protocol, or a specific disease cohort, while relying on extensive expert annotations. Consequently, substantial model redevelopment and additional annotation may be required when the anatomy, imaging characteristics, or disease distribution changes. This dissertation develops and validates a series of representation learning methods that use diverse CT images and radiology reports to build reusable models for CT image analysis. The first part develops anatomy-aware self-supervised methods for CT-based organ and lesion segmentation. SS-UNet leverages masked image modeling with sparse convolutional networks to learn anatomical representations from unlabeled CT volumes. The resulting representations improve robustness and transferability across organ and lesion segmentation tasks in CT, magnetic resonance imaging, and positron emission tomography compared to existing baselines. AnatoMask extends this framework through reconstruction-guided self-masking, further improving segmentation under limited labeled-data settings. The second part develops OpenVocabCT, a vision--language model that aligns CT volumes with both complete radiology reports and organ-level descriptions. This approach enables text-driven segmentation of organs and tumors and improves generalization to varied and previously unseen clinical prompts across public and institutional datasets. The third part extends language supervision to disease-oriented CT interpretation. MedVista3D jointly aligns localized regions and whole CT volumes with semantically enriched radiology text, supporting zero-shot disease classification, medical visual question answering, automated reporting, and transfer to downstream CT imaging tasks. The final part develops FlexiCT, a family of CT foundation models trained through stage-wise pretraining on 266,227 CT volumes from 56 public datasets. FlexiCT supports segmentation, registration, disease classification, tumor-phenotype retrieval, and vision--language analysis within a common model family. To evaluate clinical specialization, DINO-CardiacSeg adapts CT foundation-model to contrast-enhanced radiotherapy simulation CT for segmenting 21 cardiac substructures for dosimetry and cardiotoxicity assessment. Together, these studies develop a progression of representation learning frameworks for radiology and radiation oncology. Image-based self-supervision provides reusable spatial features. Radiology reports introduce flexible clinical concepts and disease semantics. Targeted adaptation extends general CT representations to specialized clinical applications. By reducing dependence on manual annotation and improving robustness across heterogeneous imaging settings, these methods support more scalable CT image analysis, including organ and tumor delineation for treatment planning and disease assessment in diagnostic radiology.