Geospatial Data
Introduction
Geospatial data in the areas of socioeconomic status, air quality, and a variety of agricultural and meteorological exposures have been linked to ~149k addresses for ~99k PLCO participants. This page provides a brief explanation for how address histories were obtained, how addresses were geocoded/linked to exposures and considerations when selecting which address to use. Data cleaning to eliminate small cells for distribution are also described.
Addresses and Geocoding
During the PLCO trial only the current address of a participant was retained for administrative purposes. In 2017, an effort to search participants in Lexis Nexis (LN) was conducted in attempt to construct a participant’s full address history. Only participants who were deceased or had agreed to participate in extended follow-up, transferred to the Central Data Collection Center (CDCC) and were not from the Marshfield center were searched.
Lexis Nexis addresses found were deduplicated, adjusted to dob/dod and were aligned with the current address on file so that there were no gaps or overlaps in the addresses available for analysis. Generally, the first time an address is found in the LN search, the time at that address is started until a new address is found. Note, <3% of the addresses were not found in LN and the current address is the only source, for these timing was set to start at randomization.
The CDCC geocoded each address a geographical identifier for latitude and longitude using their Geographic Information System (GIS). Levels of precision were identified and addresses available for analysis were limited to a zip code level or closer. Geocodes were used to link several georeferenced databases to each address to obtain a variety of census and environmental data.
Due to the uncertainty of the exact timing at each address, there are limitations to this database.
Address Selection
Roughly a third of the ~99k participants with geocoded data have more than one address available for analysis. Standard data is available for request for the address at entry (trial randomization) or exit (first cancer or mortality). The number of days from start of address to event is included to represent a minimum known time at address prior to event.
Sometimes, events occur prior to the first address we have for a participant. In these situations, the first address after the event is used and the number of days from start to event is negative.
Exposure Data Cleaning
Most of the linked exposure data is continuous and comes from public geospatial databases. To decrease any potential participant identification from those databases, new versions of the data were derived so that no values have less than 10 PLCO addresses. The following steps were completed when necessary to remove small cells with the least manipulation possible.
- Values rounded to less precision.
- Outliers grouped and assigned to the median of those outliers. The range of the outliers is documented in the data dictionary.
Example: 8,103 unique greenspace values were received ranging from -0.1996 to 0.9654. After rounding to nearest .01 and grouping the outliers we now have 104 values. The small cells on the lower end ranging from -0.1996 to -0.1150 were assigned to the value median -0.16. The small cells on the upper end ranging from 0.8500 to 0.9654 were assigned to the median 0.86.
Census 1990
Census 2000
Census 2010
Urbanicity 1990, 2000, 2010
- Indicator of whether a participant falls within a Census 2000 Urban Area (UA) boundary
- United States Census Bureau Urban and Rural Geographic Areas
Yost Score 2011-2018
- SES Index based on education, income, poverty rate, housing, rent, unemployment rate and mix
- Socioeconomic status and breast cancer incidence in California for different race/ethnic groups
ADI Ranks 2013, 2015, 2018
- Area Deprivation Index – National Rank
- About the Neighborhood Atlas® and Area Deprivation Index (ADI)
| Census 1990 | Census 2000 | Census 2010 | ACS (2010-2012) |
Urbanicity (1990/2000/2010) |
Yost Index (2011-2018) |
Area Deprivation Index (2013/2015/2018) |
|
|---|---|---|---|---|---|---|---|
| SES | X | X | |||||
| Urbanicity | X | ||||||
| Median Household Income | X | X | |||||
| % Female Unemployed | X | ||||||
| % Male Unemployed | X | X | X | ||||
| % Total Unemployed | X | X | X | ||||
| % Hispanic/Black/White | X | X | X | X | |||
| % Not in Labor Force | X | X | X | ||||
| % Residents 65+ | X | X | X | X | |||
| Total Population | X | X | X | X | |||
| Average Household Size | X | X | |||||
| Mortgage Status | X | ||||||
| Median Home Value | X | X | |||||
| % Below Poverty Line | X | X | |||||
| % Receiving Public Assistance | X | X | |||||
| % Crowding | X | X | |||||
| % Female Head of Household | X | X | X | ||||
| % Renters | X | X | X | ||||
| % Vacant Homes | X | X | X | ||||
| Broadband | X | ||||||
| Sex by industry | X | ||||||
| Gini Index | X | ||||||
| Vehicle Ownership | X | ||||||
| Population Education | X | ||||||
| Population Health Insurance | X | ||||||
| Tenure by Occupants | X | ||||||
| Tenure by Telephone | X | ||||||
| Industry Specific Jobs | X |
PM2.5 and Components
Gridded estimates (0.01 by -.01 degree) of ground-level fine particulate matter (PM2.5) total and compositional mass concentrations over North America through combinations of Aerosol Optical Depth (AOD) retrievals from the NASA MODIS, MISR, and Sea WIFS instruments with the GEOS-Chem chemical transport model, and subsequently calibrated to regional ground-based observations of both total and compositional mass using Geographically Weighted Regression (GWR).
van Donkelaar A, Martin RV, Li C, Burnett RT. Regional Estimates of Chemical Composition of Fine Particulate Matter Using a Combined Geoscience-Statistical Method with Information from Satellites, Models, and Monitors. Environ Sci Technol. 2019 Mar 5;53(5):2595-2611. doi: 10.1021/acs.est.8b06392. Epub 2019 Feb 12. PMID: 30698001.
The following variables are included:
Ground-level fine particulate matter (PM2.5) total (air_pm25_2000-2018) -
Component concentrations (percents):
- Ammonia (air_nh4_2000-2017)
- Black Carbon (air_bc_2000_2017)
- Nitrate (air_no3_2000-2017)
- Organic Matter (air_om_2000-2017)
- Sulfate (air_so4_2000-2017)
- Soil (air_soil_2000-2017)
- Sea Salt (air_ss_2000-2017)
Nitrogen Dioxide (NO2)
Nitrogen Dioxide (NO2) data includes NO2 estimates available at centroid of census blocks as well as averages for block group and census tract. Derived from land-use regression model by Bechle et al (2015)*.
Annual values for 2000-2010 are available and were created by averaging the monthly LUR tract level Model Estimates. 2000 US Census is the boundary file used. Data does not include Alaska or Hawaii.
Bechle MJ, Millet DB, Marshall JD. National Spatiotemporal Exposure Surface for NO2: Monthly Scaling of a Satellite-Derived Land-Use Regression, 2000-2010. Environ Sci Technol. 2015 Oct 20;49(20):12297-305. doi: 10.1021/acs.est.5b02882. Epub 2015 Oct 9. PMID: 26397123.
Ozone
Ozone data comes from the US Environmental Protection Agency. These are census-tract level, annual estimates from 2002-2017.
The following are included for each year:
- Minimum observation
- 5th percentile
- 25th percentile
- Mean
- 75th percentile
- 95th percentile
- Maximum
- Median
- Standard deviation
EPA's Downscaling Output files
A Bayesian space-time downscaler model is used to "fuse" daily ozone (8-hr max) and fine particulate air (24-hr average) monitoring data from the National Air Monitoring Stations/State and Local Air Monitoring Stations (NAMS/SLAMS) with 12 km gridded output from the Models-3/Community Multiscale Air Quality (CMAQ) model.
Daily predictions were estimated at the 2010 US Census Tract centroid locations. The 2002-2008 downscaling outputs were run on 2002-2008 inputs using 2010 Census Tracts to be comparable across the years with the post-2018 downscaler output. Daily predictions were averaged into monthly and annual values.
Land Cover Data
Land cover data from the USGS National Land Cover Database (NLCD) includes types of land cover in concentric buffers around the address at fixed distances of 100, 250, 500, 750, and 1000m. Land cover data for 2001, 2004, 2006, 2008, 2011, 2013, and 2016 were obtained from the Multi-Resolution Land Characteristics Consortium's (MRLCs) National Land Cover Database (NLCD) products during April, 2020. Datasets have a 30-meter resolution and cover the contiguous U.S. (Alaska and Hawaii are excluded).
Each year of land cover data contains the same land cover categories. Of the 20 land cover categories defined by MRLC, four are only available in Alaska and therefore are not represented herein. These four are: 51 - Dwarf Scrub, 72 - Sedge/Herbaceous, 73 - Lichens, and 74 - Moss. The remaining 16 categories (plus one "unclassified" category) which apply herein are listed below. More detailed definitions of each category can be found on the MRLC website.
Value 0 - Unclassified (Note: this field can be used to assess the completeness of NLCD data within the buffer)
Value 11 - Open water
Value 12 - Perennial ice/snow
Value 21 - Developed, open space
Value 22 - Developed, low intensity
Value 23 - Developed, medium intensity
Value 24 - Developed, high intensity
Value 31 - Barren land
Value 41 - Deciduous forest
Value 42 - Evergreen forest
Value 43 - Mixed forest
Value 52 - Shrub/scrub
Value 71 - Grassland/herbaceous
Value 81 - Pasture/hay
Value 82 - Cultivate crops
Value 90 - Woody wetlands
Value 95 - Emergent herbaceous wetlands
Elevation
Elevation raster data at 800-meter resolution for the contiguous U.S. were obtained from PRISM Climate Group.
Soil pH
Gridded National Soil Survey Geographic Database (gNATSGO) data were downloaded from the USDA NRCS Dept of Agriculture on 4/21/20. A 30-meter soil dataset representing "pH (1-to-1 Water)" was generated using soil depth set to 0-10cm and an aggregation method of "weighted average."
Greenspace
MODIS NDVI (Normalized Difference Vegetation Index) is estimated at a 250-meter resolution for years 2000-2014. NDVI data represent the measure of vegetation/greenness ranging from -1 (water) to 1 (complete vegetation) on a given date. NDVI variable names have format "nYYMMDD" and are defined as follows:
n stands for NDVI
YY is the two digit year (e.g., 00 is 2000)
MM is the two digit month (e.g., 01 is January)
DD is the two digit day of the month
(Example: n000406 is NDVI from April 6, 2000)
Noise
Noise data were obtained from a georeferenced noise model of expected environmental sound level. These acoustical data are estimated from 1.5 million hours of long-term measurements from 492 urban and rural sites located across the contiguous United States during 2000-2014. The resulting non-time-varying (integrated over 2000-2014) geospatial sound model enabled mapping of sound levels at 270m resolution.
Each variable of noise data represents a measure of anthropogenic sound (_ant), natural sound (_nat) and total existing sound (_exi):
L10d_ant
L10d_exi
L10d_nat
L50d_ant
L50d_ant_nt
L50d_exi
L50d_nat
L90d_ant
L90d_exi
L90d_nat
Ldn_exi
Leq24_exi
Lnight_exi
Light-at-Night (LAN)
Light-at-Night (LAN) rasters were sourced from personal correspondence with Peter James (HSPH, Harvard - pjames@hsph.harvard.edu) on 5/15/18 and represent annual composites of LAN. P. James processed raw data from satellite files and calibrated radiance across different satellites so that raster files were comparable. Raster values range from 0-3,984.69, depending on the dataset. Technically, the rasters are unit-less due to the lack of on-board calibration systems, but P. James converted them to approximate w/cm2/sr units. The conversion to radiance may not be exact. To read more about P. James' calibration methods visit MDPI. Data were projected by A. Flory into USA Contiguous Lambert Conformal Conic (WKID: 102004). Data have ~1.26 km resolution.
Each year-variable of LAN data represents data from the following date ranges:
LAN_1996:Mar16,1996–Feb12,1997
LAN_1999:Jan19,1999–Dec11,1999
LAN_2000:Jan3,2000–Dec29,2000
LAN_2003:Dec30,2002–No27,2003
LAN_2004:Jan18,2004–Dec16,2004
LAN_2006:Nov28,2005–Dec24,2006
LAN_2010: Jan 11, 2010 – Dec 9, 2010
Ultraviolet-B Light Data
Measures reflect monthly measures of UVB (Ultraviolet-B light) on a 1x1 degree grid – long-term average (1982-1992) from images captured from the NASA TOMS satellite. Coordinates are standard latitude and longitude and reflect the midpoint of gridded cells.
Precipitation, Temperature, Wind Data
A variety of meteorological variables including precipitation, temperature, wind direction and wind speed were obtained from the University of Oregon’s PRISM Weather database. Information on PRISM can be found on their website.
Average monthly values for precipitation, minimum temperature, and maximum temperature for years 2000-2003 are included. Temperature data are in degrees Celsius (-100 to 100) and precipitation data are in millimeters (0-15,000).
- Monthly prevailing wind direction for each year, 1979-2021 (NARR)