Artificial Intelligence (AI)-Assisted Health Informatics to Stratify Population for Early Screening of Cancer (Sibling Project to PLCO-2067)
Principal Investigator
Name
Gregory Hart
Degrees
Ph.D.
Institution
University of Guam
Position Title
Assistant Professor
About this CDAS Project
Study
PLCO
(Learn more about this study)
Project ID
PLCO-2078
Initial CDAS Request Approval
Aug 31, 2026
Title
Artificial Intelligence (AI)-Assisted Health Informatics to Stratify Population for Early Screening of Cancer (Sibling Project to PLCO-2067)
Summary
Electronic Medical Records (EMRs) are designed to gather and store health information. We believe that this data could be mined to further improve cancer care. Our goal is to predict individual cancer risk from mined health informatics from EMRs using artificial intelligence (AI) tools. This cancer risk-based population stratification will allow resources to be concentrated on those most in danger and allowing those with low risk to avoid screenings that can be expensive, comfortable, and have their own risks. Our aims for the next two years are to move toward this vision. First, we will develop a multi-input/output deep neural network to predict cancer risk through common health informatics e.g., the National Health Interview Survey from 1997-2022 and the prostate, lung, colorectal, and ovarian (PLCO) trials data. Once we show that simple health data can be used to predict cancer risk and stratify patients, we aim to show that this information can be extracted from EMRs for real-time updates to detect cancer in early stage. Therefore, our 2nd aim is to develop a deep reinforcement learning (DRL) tool for predicting those same risks based on real-time inputs of health informatics. Our 3rd aim is to test our AI tool in clinical environment and to test it for two-way data sharing to help populations with high-risk for cancer by lowering their risk, promoting early screening, allowing for symptom tracking, and communicating their data with healthcare providers. We will establish cancer scoring indices (CSIs) for each individual cancer type e.g., a high CSI for a particular cancer type means a high risk for that cancer type accordingly. Completion of the project would assist physicians in determining cancer risks and needed screenings through CSI. The tool may be integrated into an EMR system and be available on websites and portable devices, making it immediately available for clinicians to predict cancer and improve patient outcomes by detecting cancer at an early stage.
Aims
Aim 1: Develop multi-input/multi-output deep neural network (DNN) for risk prediction of cancer. Our first aim is to develop a model that can predict a person’s risk for cancer from simple health informatics data such as age, race, family history, BMI, tobacco use, blood pressure, etc. This model will be designed robustly allowing for predictions to be made even if a few of the inputs are missing. It offers a convenient and cost-effective way to identify the risk, lowering the burden of cancer as earlier detection and interventions become possible. For this purpose, we will use two large medical datasets (1) the National Health Interview Survey (NHIS) from 1997 -2022 and (2) the prostate, lung, colorectal, and ovarian (PLCO) trails.
Aim 2: Extend deep neural network with reinforcement learning to make predictions based on real-time inputs of health informatics. Our second aim is to expand the model from aim 1, allowing the use of real-time inputs with implementation of reinforcement learning and will build a deep reinforcement learning (DRL) model. The model in aim 1 is useful, but it is a static prediction based on data about one’s health at a specific point in time. We want our model to be able to update its predictions as one’s health information changes. Also, there is information on the changes that happen over time. Being able to look at the trajectory of one’s health and not just a single point will allow the model to further learn and improve its predictions. So, we will be using the UK Biobank database where half a million participants from 2006 - 2010 were enrolled. In this dataset many types of follow-up and additions are frequently made best suited to dynamic prediction.
Collaborators
Gregory Hart University of Guam
Danielle Balmores University of Guam
Wazir Muhammad Florida Atlantic University