Skip to Main Content
An official website of the United States government

Developing scalable machine learning-based approaches to classifying significant incidental findings in patients undergoing lung cancer screening

Principal Investigator

Name
Ilana Gareen

Institution
Brown University

Position Title
Professor

Email
ilana_gareen@brown.edu

About this CDAS Project

Study
NLST (Learn more about this study)

Project ID
NLST-1517

Initial CDAS Request Approval
Jul 20, 2026

Title
Developing scalable machine learning-based approaches to classifying significant incidental findings in patients undergoing lung cancer screening

Summary
This project will evaluate scalable natural language processing (NLP) approaches for identifying significant incidental findings (SIFs) from free-text comments describing "other significant abnormalities" in patients receiving low-dose CT as part of the National Lung Screening Trial (NLST). SIFs identified during lung cancer screening may have important clinical implications and can lead to additional diagnostic evaluation or follow-up care. However, identifying these findings from radiology reports currently relies heavily on manual review or rule-based methods, both of which can be labor-intensive and difficult to scale to large datasets. The proposed work will investigate whether modern machine learning and artificial intelligence methods can accurately classify SIFs directly from free text. Specifically, we will evaluate machine learning and transformer-based NLP models applied to free text comments describing SIFs detected in the NLST. The project will compare the performance of alternative approaches to existing rule-based methods and assess their potential to support more efficient large-scale analysis of radiology report data. All analyses involving NLST data will be conducted within secure local computing infrastructure using locally hosted models. Findings from this work may help support future large-scale studies of SIFs and demonstrate approaches for applying modern NLP methods to clinically relevant free-text data in large screening studies.

Aims

Incidental findings identified during lung cancer screening may have important clinical implications, but identification of these findings from radiology reports often relies on labor-intensive manual review or rule-based methods that can be difficult to scale. This project will evaluate the feasibility of applying modern natural language processing (NLP) methods to free-text comments describing “other significant abnormalities” in the National Lung Screening Trial (NLST). The overall goal is to assess whether machine learning and transformer-based NLP models can provide an accurate and scalable approach for identifying significant incidental findings (SIFs) from radiology report text.

Specific aims are to:

• Evaluate scalable NLP approaches for identifying significant incidental findings (SIFs) from free-text comments describing “other significant abnormalities” in the NLST.
• Compare the performance of machine learning and transformer-based NLP models to existing rule-based approaches using manually reviewed SIF classifications as the reference standard.
• Assess the ability of NLP models to identify SIFs across varying levels of clinical ambiguity and linguistic complexity.
• Generate preliminary evidence regarding the feasibility of applying modern NLP methods to large-scale text-based classification tasks in cancer screening studies.

Results from this work may help support future large-scale studies using radiology report text and provide insight into the potential role of modern NLP methods in cancer screening research.

Collaborators

Ilana Gareen Brown University
Yifei Feng Brown University