PLCO Germline Materials for Integrated Whole-Genome Discovery of Lung Cancer Susceptibility with the Sherlock-Lung Study
Principal Investigator
Name
Maria Teresa Landi
Degrees
M.D., Ph.D.
Institution
National Cancer Institute, National Institutes of Health
Position Title
Senior Investigator & Senior Advisor for Genetic Epidemiology
Email
landim@nih.gov
About this CDAS Project
Study
PLCO
(Learn more about this study)
Project ID
PLCO-2016
Initial CDAS Request Approval
Jan 15, 2026
Title
PLCO Germline Materials for Integrated Whole-Genome Discovery of Lung Cancer Susceptibility with the Sherlock-Lung Study
Summary
Lung cancer remains a leading cause of cancer mortality, and inherited susceptibility is incompletely characterized—particularly for rare variation, structural variation, and non-coding regulatory regions that are not well captured by genotyping arrays alone. The PLCO Cancer Screening Trial biorepository offers an exceptional, deeply annotated population-based resource with linked demographic and clinical data.
We propose to use PLCO germline materials from lung cancer cases and a set of non-cancer controls and integrate them with the Sherlock-Lung cohort (~5,300 whole-genome sequencing [WGS] samples) and additional datasets, comprising over 10,000 individuals that we are actively analyzing in collaboration with other DCEG investigators (e.g., Wendy Wong) and Cancer Genomics Research Laboratory (CGR), to enable comprehensive germline analyses spanning coding and non-coding regions. Specifically, we are requesting germline material from all primary lung cancer cases (~3100) and a set of 1,000 controls frequency matched to cases by ancestry, age (+/- 5 years), sex, and tobacco smoking status.
Importantly, we will leverage genotyping data that has already been generated by PLCO for ancestry inference and linkage to existing GWAS resources. Our primary laboratory activity will be WGS (possibly long-read WGS), and we have obtained special funding to cover the sequencing cost of these samples. WGS will enable discovery of common and rare single-nucleotide variants, indels, structural variants, and non-coding regulatory variants associated with lung cancer susceptibility. These data will also be useful to other investigators using PLCO data. We will meta-analyze these data with Sherlock-Lung WGS and other datasets (23andMe, UKBB, AllofUs, 1000Genome, Genomics England and others) reaching over 10,000 lung cancer cases to improve discovery power and facilitate replication.
A key goal is to combine WGS-based discovery with existing GWAS-derived polygenic risk scores (PRS) to achieve a more comprehensive understanding of genetic susceptibility—integrating common and rare, coding and non-coding contributions to lung cancer risk.
We will need demographic and clinical data for the analysis, including ancestry, age, sex, tobacco smoking status, lung cancer histology, PLCO arm status, and possibly subjects' residence and other variables.
Aims
Aim 1: Generate WGS data from PLCO germline specimens of ~3100 lung cancer cases and ~1,000 non-cancer controls) to discover coding and non-coding susceptibility variants. Perform WGS on PLCO germline materials to identify common and rare SNVs/indels, structural variants, and non-coding regulatory variants associated with lung cancer susceptibility. Apply standardized QC and harmonized processing to ensure robust downstream analyses and cross-cohort integration.
Aim 2: Meta-analyze PLCO WGS with Sherlock-Lung (~5,300 WGS) and other large datasets (e.g., UK Biobank, All of Us) totaling >10,000 case samples to improve discovery, replication, and generalizability of inherited risk signals. Where possible we will use similar pipelines for analyses and annotation, and prioritize high-confidence coding and non-coding susceptibility variants and assess consistency across cohorts and ancestries (where available).
Aim 3: Combine WGS-based findings with existing PLCO genotyping/GWAS resources and GWAS-derived PRS to model lung cancer susceptibility across common/rare and coding/non-coding genetic architectures. Leverage existing PLCO genotyping and PRS to quantify how WGS-identified rare and structural variants contribute to risk beyond common-variant polygenic burden. Develop integrated analytic frameworks that jointly evaluate PRS and WGS-derived rare variant effects to refine biological insight and improve risk stratification.
Collaborators
Maria Teresa Landi National Cancer Institute, National Institutes of Health
Huu Phuc Hoang National Cancer Institute, National Institutes of Health
John McElderry National Cancer Institute, National Institutes of Health
Xuan Li National Cancer Institute, National Institutes of Health
Shuk Wan Wendy Wong National Cancer Institute, National Institutes of Health
Ben Jordan National Cancer Institute, National Institutes of Health
Komal Jain National Cancer Institute, National Institutes of Health
Wei Zhu National Cancer Institute, National Institutes of Health