Developing scalable machine learning-based approaches to classifying significant incidental findings in patients undergoing lung cancer screening
Principal Investigator
Name
Ilana Gareen
Institution
Brown University
Position Title
Professor
Email
ilana_gareen@brown.edu
About this CDAS Project
Study
NLST
(Learn more about this study)
Project ID
NLST-1517
Initial CDAS Request Approval
Jul 20, 2026
Title
Developing scalable machine learning-based approaches to classifying significant incidental findings in patients undergoing lung cancer screening
Summary
This project will evaluate scalable natural language processing (NLP) approaches for identifying significant incidental findings (SIFs) from free-text comments describing "other significant abnormalities" in patients receiving low-dose CT as part of the National Lung Screening Trial (NLST). SIFs identified during lung cancer screening may have important clinical implications and can lead to additional diagnostic evaluation or follow-up care. However, identifying these findings from radiology reports currently relies heavily on manual review or rule-based methods, both of which can be labor-intensive and difficult to scale to large datasets. The proposed work will investigate whether modern machine learning and artificial intelligence methods can accurately classify SIFs directly from free text. Specifically, we will evaluate machine learning and transformer-based NLP models applied to free text comments describing SIFs detected in the NLST. The project will compare the performance of alternative approaches to existing rule-based methods and assess their potential to support more efficient large-scale analysis of radiology report data. All analyses involving NLST data will be conducted within secure local computing infrastructure using locally hosted models. Findings from this work may help support future large-scale studies of SIFs and demonstrate approaches for applying modern NLP methods to clinically relevant free-text data in large screening studies.
Aims
Incidental findings identified during lung cancer screening may have important clinical implications, but identification of these findings from radiology reports often relies on labor-intensive manual review or rule-based methods that can be difficult to scale. This project will evaluate the feasibility of applying modern natural language processing (NLP) methods to free-text comments describing “other significant abnormalities” in the National Lung Screening Trial (NLST). The overall goal is to assess whether machine learning and transformer-based NLP models can provide an accurate and scalable approach for identifying significant incidental findings (SIFs) from radiology report text.
Specific aims are to:
• Evaluate scalable NLP approaches for identifying significant incidental findings (SIFs) from free-text comments describing “other significant abnormalities” in the NLST.
• Compare the performance of machine learning and transformer-based NLP models to existing rule-based approaches using manually reviewed SIF classifications as the reference standard.
• Assess the ability of NLP models to identify SIFs across varying levels of clinical ambiguity and linguistic complexity.
• Generate preliminary evidence regarding the feasibility of applying modern NLP methods to large-scale text-based classification tasks in cancer screening studies.
Results from this work may help support future large-scale studies using radiology report text and provide insight into the potential role of modern NLP methods in cancer screening research.
Collaborators
Ilana Gareen Brown University
Yifei Feng Brown University