AEGIS-Lung — Deep-Learning Development and Internal Validation of an Integrated Lung-Cancer Risk Model in the NLST Cohort
Principal Investigator
Name
Aneel Paulus
Degrees
M.D., M.S.
Institution
West Eastern Health
Position Title
CEO
About this CDAS Project
Study
NLST
(Learn more about this study)
Project ID
NLST-1522
Initial CDAS Request Approval
Jul 14, 2026
Title
AEGIS-Lung — Deep-Learning Development and Internal Validation of an Integrated Lung-Cancer Risk Model in the NLST Cohort
Summary
Background and Rationale.
West Eastern Health is a clinical and research center whose established practice centers on the wellbeing of cancer patients. Because our patients so often reach us after late-stage diagnosis, we have expanded with our partners into preventive medicine- developing an integrated risk stratification model that fuses biopsychosocial history, laboratory markers, hereditary and
genomic factors, and imaging into a per-cancer risk profile. Lung cancer is the leading cause of cancer death, yet screening reaches few eligible high-risk individuals and generates many false positive findings. The National Lung Screening Trial (NLST)- a randomized comparison of low-dose CT (LDCT) versus chest radiography in 53,454 high-risk current and former smokers- pairs smoking and clinical histories, screen-detected nodule data, LDCT imaging, and adjudicated lung-cancer incidence, histology, stage, and mortality. This makes it an ideal resource for training and internally validating our model’s lung-specific module.
Objectives.
1. Develop and internally validate deep-learning models for incident lung cancer and lung cancer mortality from integrated smoking, clinical, and screen-derived features.
2. Quantify incremental value of multi-domain and imaging-derived predictors over smoking-based models.
3. Improve nodule-level stratification to reduce false-positive referrals, and assess subgroup performance.
Study Design.
Retrospective cohort study using NLST as a labeled training corpus. Baseline demographics, smoking history (pack-years, intensity, time since cessation), medical and family history, and screen-detected nodule characteristics serve as structured inputs; LDCT-derived imaging features are incorporated where available. Incident lung cancer, histology/stage, and mortality, structured as time-to-event labels, serve as training targets. The cohort is partitioned into
training, validation, and held-out test sets, with screening arm (LDCT vs chest radiography) retained for stratified evaluation.
Analysis Plan.
We will train time-to-event deep neural networks (competing-risk-aware survival architectures) and nodule-level malignancy classifiers, tuned by cross-validation and benchmarked against established lung-cancer risk models (PLCOm2012, Bach) and nodule models (Brock/PanCan, Lung-RADS). Discrimination (time-dependent AUC / Harrell’s C-index) and calibration
(calibration curves, observed-to-expected ratios) will be assessed on held-out data, with recalibration where indicated. Incremental value will be tested via ΔC-index, net reclassification improvement, and integrated discrimination improvement; decision-curve analysis will quantify potential reduction in false-positive referrals. Subgroup analyses (age, sex, race/ethnicity, smoking status, screening arm) will assess fairness and calibration drift, acknowledging NLST’s limited diversity; missing data handled with multiple imputation; competing mortality with Fine–Gray and competing-risk models. De-identified data only; no re-identification will be attempted.
Expected Outcomes and Significance.
This work will show whether an integrated model improves lung-cancer risk prediction and nodule stratification beyond smoking-based tools, and calibrate it for safe, equitable use. The objective is to develop and internally validate these models so that, in future preventive-care use, high-risk individuals can be prioritized for screening and for smoking-cessation and
behavioral support- the psychological care that is West Eastern Health’s core expertise- while reducing false-positive workups.
Aims
Aim 1: Assemble and harmonize an ML-ready training corpus.
• Map NLST demographics, detailed smoking history (pack-years, intensity, duration, time since cessation), medical and family history, and screen-detected abnormality/nodule characteristics to the model’s multi-domain input schema; incorporate LDCT-derived imaging features where available.
• Define time-to-event labels for incident lung cancer (with histology and stage) and lung cancer mortality.
• Engineer features, encode missingness explicitly, and partition participants into training, validation, and held-out test sets with no cross-partition leakage; retain screening arm (LDCT vs chest radiography) for stratified evaluation.
Aim 2: Train deep-learning lung-cancer risk and nodule models.
• Develop time-to-event deep neural networks (DeepSurv/DeepHit-style, competing-risk aware) for incident lung cancer and mortality, and nodule-level malignancy classifiers from screen-detected features.
• Optimize via cross-validation; tune regularization and class-imbalance handling for rare event rates.
• Benchmark against established lung-cancer risk models (PLCOm2012, Bach) and nodule-malignancy models (Brock/PanCan, Lung-RADS categories).
Aim 3: Evaluate discrimination, calibration, and clinical utility.
• Assess discrimination via time-dependent AUC / Harrell’s C-index and calibration (curves, observed-to-expected ratios) on held-out data, with recalibration (isotonic/Platt) where indicated.
• Use decision-curve analysis to quantify potential reduction in false-positive referrals relative to Lung-RADS.
• Examine performance across age, sex, race/ethnicity, smoking status, and screening arm to detect bias and calibration drift, acknowledging NLST’s limited racial/ethnic diversity.
Aim 4: Quantify incremental value and produce a deployment-ready
specification.
• Quantify gains from multi-domain and imaging-derived features over smoking-based models (ΔC-index, net reclassification improvement, integrated discrimination improvement).
• Run sensitivity analyses for competing mortality (Fine–Gray and competing-risk deep learning models), screen-detection effects, and models excluding imaging features.
• Deliver a calibrated, documented lung-cancer module to inform West Eastern Health’s preventive-care risk-stratification workflow, using de-identified data only with no attempted re-identification.
Alignment note. These aims operationalize the Project Summary: Aim 1 builds the training corpus described in the study design; Aims 2–3 execute the deep-learning training and evaluation analysis plan; Aim 4 delivers the calibrated lung-cancer module that advances West mission of reducing late lung-cancer diagnosis.
Collaborators
Aneel Paulus West Eastern Health