Vision–Language Learning for Generalizable Pulmonary Nodule Characterization in Low-Dose CT Screening
Principal Investigator
Name
Benjamin Wild
Degrees
Ph.D.
Institution
Berlin Institute of Health at Charité
Position Title
Group Leader
About this CDAS Project
Study
NLST
(Learn more about this study)
Project ID
NLST-1521
Initial CDAS Request Approval
Jul 14, 2026
Title
Vision–Language Learning for Generalizable Pulmonary Nodule Characterization in Low-Dose CT Screening
Summary
Accurate detection and characterization of pulmonary nodules on low-dose CT remain challenging, limiting the widespread clinical adoption of computer-assisted decision support systems. Recent advances in multimodal vision-language learning have demonstrated considerable promise but remain insufficiently explored for low-dose CT lung nodule analysis. In our project we aim to develop and evaluate a multimodal vision-language model that learns transferable representations for lung nodule analysis. Recent studies suggest that multimodal supervision using semantic descriptors can improve model interpretability compared to conventional image-only deep learning methods. Access to NLST is essential because its unique combination of large-scale low-dose screening CT, longitudinal follow-up, and clinically confirmed outcomes cannot be replicated by existing public datasets such as LIDC-IDRI. NLST provides a uniquely valuable, large-scale longitudinal screening cohort that broadens population diversity while enabling rigorous development and evaluation of generalizable models. Together with complementary datasets, NLST will be used to train models using the expert-derived semantic descriptors (e.g., spiculation, lobulation, texture, and calcification) linked to annotated pulmonary nodules in CT images.
Rather than relying solely on CT imaging, the proposed framework will jointly learn from image data and expert-derived semantic descriptors using a CLIP-inspired vision-language architecture and work as multimodal representation learning framework for lung nodule detection and automated classification. Semantic descriptors capture clinically meaningful information that is difficult to infer directly from image appearance alone and therefore provide complementary supervision. Using both modalities improves model robustness, interpretability, and generalization across institutions. Besides malignancy prediction, our aim is to validate the proposed model by comparing it with current state-of-the-art methods across multiple downstream tasks, including malignancy prediction, invasiveness assessment, and pathology-related prediction. The availability of longitudinal screening examinations and confirmed clinical outcomes enables evaluation of temporal disease progression and assessment of clinically relevant endpoints beyond single-timepoint malignancy prediction. Ultimately, this work aims to improve computer-assisted characterization of pulmonary nodules, facilitating earlier identification of high-risk lesions while reducing unnecessary diagnostic procedures and follow-up examinations.
Model development will leverage the complementary strengths of different datasets. LIDC-IDRI provides expert semantic annotations, NLST contributes large-scale screening CT examinations with longitudinal clinical outcomes, and an institutional dataset will be used for additional external validation. In contrast to many other existing methods our goal is to extend prediction capability, including malignancy prediction, invasiveness assessment, and selected pathology-related endpoints. Model performance will be compared with current state-of-the-art approaches through both internal, and independent external validation to assess robustness and generalizability. Performance will be evaluated using AUROC, AUPRC, sensitivity, specificity, calibration, and clinically relevant subgroup analyses. Ablation experiments will quantify the contribution of semantic information compared with image-only models. All experiments will follow reproducible training and evaluation protocols with predefined data splits and external validation cohorts.
Aims
- Develop a multimodal vision-language representation learning framework that jointly learns from low-dose CT images and expert-derived semantic descriptors of pulmonary nodules.
- Leverage the complementary strengths of the NLST, LIDC-IDRI, and an institutional dataset to train and validate robust, generalizable models for pulmonary nodule analysis.
- Evaluate whether multimodal representation learning improves performance over image-only state-of-the-art approaches across clinically relevant downstream tasks, including malignancy prediction, invasiveness assessment, and selected pathology-related endpoints.
- Assess model robustness, calibration, interpretability, and generalizability through internal validation, independent external validation, and ablation studies quantifying the contribution of semantic information.
- Generate transferable multimodal representations that can support future computer-assisted analysis of pulmonary nodules and related low-dose CT screening applications.
Collaborators
Benjamin Wild Berlin Institute of Health at Charité
Georg von Arnim Berlin Institute of Health at Charité