[Paper Review] A County-level Dataset for Informing the United States' Response to COVID-19
A machine-readable, county-level dataset aggregating COVID-19 time-series, NPIs, mobility, and 300+ demographic/socioeconomic variables to study regional spread and inform rollback of interventions; code and data are available publicly.
As the coronavirus disease 2019 (COVID-19) continues to be a global pandemic, policy makers have enacted and reversed non-pharmaceutical interventions with various levels of restrictions to limit its spread. Data driven approaches that analyze temporal characteristics of the pandemic and its dependence on regional conditions might supply information to support the implementation of mitigation and suppression strategies. To facilitate research in this direction on the example of the United States, we present a machine-readable dataset that aggregates relevant data from governmental, journalistic, and academic sources on the U.S. county level. In addition to county-level time-series data from the JHU CSSE COVID-19 Dashboard, our dataset contains more than 300 variables that summarize population estimates, demographics, ethnicity, housing, education, employment and income, climate, transit scores, and healthcare system-related metrics. Furthermore, we present aggregated out-of-home activity information for various points of interest for each county, including grocery stores and hospitals, summarizing data from SafeGraph and Google mobility reports. We compile information from IHME, state and county-level government, and newspapers for dates of the enactment and reversal of non-pharmaceutical interventions. By collecting these data, as well as providing tools to read them, we hope to accelerate research that investigates how the disease spreads and why spread may be different across regions. Our dataset and associated code are available at github.com/JieYingWu/COVID-19_US_County-level_Summaries.
Motivation & Objective
- Motivate data-driven analysis of COVID-19 spread by county in the United States.
- Provide a machine-readable, multi-source dataset combining epidemiological, mobility, and socioeconomic variables.
- Enable analysis of how regional differences affect transmission and NPIs effectiveness.
Proposed method
- Aggregate per-county data from governmental, journalistic, and academic sources into a machine-readable CSV and accompanying data pipeline.
- Combine over 300 variables including population, demographics, housing, climate, transit, and healthcare capacity.
- Incorporate time-series data for infections/deaths and dates of NPIs and rollbacks using ordinal dates for machine-readability.
- Aggregate out-of-home activity data from SafeGraph and Google mobility reports at the county level.
- Impute missing static data with state-wide averages where appropriate.
- Provide code and repository to read and utilize the dataset for epidemiological forecasting and policy analysis.
Experimental results
Research questions
- RQ1What county-level factors (demographics, economy, climate, mobility, healthcare capacity) correlate with COVID-19 spread and severity?
- RQ2How do non-pharmaceutical interventions (NPIs) and their rollbacks at the county/state level relate to subsequent infection trends across counties?
- RQ3Can machine-learning approaches identify the most relevant factors informing effective, graduated rollback of isolation measures?
Key findings
- The dataset includes over 300 variables for 3220 county-equivalents (plus states, DC, and the US) with time-series COVID-19 data and NPI dates.
- Out-of-home activity and mobility data from SafeGraph and Google mobility reports are provided to help assess compliance and behavioral changes.
- Imputation is applied for missing static data using state-wide averages to maintain county-level coverage.
- The resource highlights how county-level contextual factors relate to disease spread and intervention effects, enabling data-driven forecasting and policy analysis.
- The repository offers machine-readable formats and tooling to accelerate epidemiological modeling and scenario analysis.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.