[논문 리뷰] A County-level Dataset for Informing the United States' Response to COVID-19
기계 판독 가능하고 카운티 단위의 데이터 세트로 COVID-19 시간 시계열, 비약물적 개입(NPIs), 모빌리티, 300개 이상의 인구사회경제 변수들을 모아 지역 확산을 연구하고 개입의 롤백을 알리려는 데이터 세트; 코드와 데이터는 공개적으로 제공됩니다.
As the coronavirus disease 2019 (COVID-19) continues to be a global pandemic, policy makers have enacted and reversed non-pharmaceutical interventions with various levels of restrictions to limit its spread. Data driven approaches that analyze temporal characteristics of the pandemic and its dependence on regional conditions might supply information to support the implementation of mitigation and suppression strategies. To facilitate research in this direction on the example of the United States, we present a machine-readable dataset that aggregates relevant data from governmental, journalistic, and academic sources on the U.S. county level. In addition to county-level time-series data from the JHU CSSE COVID-19 Dashboard, our dataset contains more than 300 variables that summarize population estimates, demographics, ethnicity, housing, education, employment and income, climate, transit scores, and healthcare system-related metrics. Furthermore, we present aggregated out-of-home activity information for various points of interest for each county, including grocery stores and hospitals, summarizing data from SafeGraph and Google mobility reports. We compile information from IHME, state and county-level government, and newspapers for dates of the enactment and reversal of non-pharmaceutical interventions. By collecting these data, as well as providing tools to read them, we hope to accelerate research that investigates how the disease spreads and why spread may be different across regions. Our dataset and associated code are available at github.com/JieYingWu/COVID-19_US_County-level_Summaries.
연구 동기 및 목표
- 미국에서 카운티별로 COVID-19 확산에 대한 데이터 기반 분석을 고취한다.
- 역학, 모빌리티, 사회경제 변수들을 결합한 기계 판독 가능한 다중 소스 데이터 세트를 제공한다.
- 지역 차이가 전파 및 NPIs 효과에 어떤 영향을 미치는지 분석할 수 있게 한다.
제안 방법
- 정부/언론/학술 자료에서 카운티별 데이터를 모아 기계 판독 가능한 CSV와 동반 데이터 파이프라인으로 만든다.
- 인구, 인구통계, 주거, 기후, 교통, 건강관리 역량 등 300개 이상 변수를 결합한다.
- 감염/사망의 시계열 데이터와 NPI 및 롤백의 날짜를 기계 판독성을 위해 서수 날짜를 사용해 포함한다.
- SafeGraph 및 Google 모빌리티 보고서의 외출 활동 데이터를 카운티 수준으로 집계한다.
- 적절한 경우 주별 평균으로 누락된 정적 데이터를 대체한다.
- 데이터 세트를 역학적 예측 및 정책 분석에 활용하기 위한 코드와 저장소를 제공한다.
실험 결과
연구 질문
- RQ1어떤 카운티 수준의 요인들(인구통계, 경제, 기후, 모빌리티, 건강 관리 역량)이 COVID-19 확산과 중증도에 상관관계가 있는가?
- RQ2카운티/주 차원의 비약물적 개입(NPIs) 및 그 롤백이 이후의 감염 추세와 각 카운티 간의 관계에 어떻게 반영되는가?
- RQ3기계학습 방법이 효과적이고 점진적인 격리 조치의 롤백에 가장 관련된 요인을 식별할 수 있는가?
주요 결과
- 이 데이터 세트는 감염과 NPI 날짜를 포함한 시간 시계열 데이터와 300개 이상의 변수로 구성된 3220개 카운티에 대응하는 상위 표현(주, DC, 미국 포함)을 포함한다.
- SafeGraph 및 Google 모빌리티 보고서의 외출 활동 및 이동성 데이터가 제공되어 준수 여부와 행동 변화 평가에 도움을 준다.
- 주별 평균을 사용한 누락된 정적 데이터의 대체를 통해 카운티 수준 커버리지를 유지한다.
- 카운티 수준의 맥락적 요인이 질병 확산 및 개입 효과와 어떻게 관련되는지 하이라이트하여 데이터 기반 예측 및 정책 분석을 가능하게 한다.
- 저장소는 역학 모델링 및 시나리오 분석을 가속화하는 기계 판독 가능한 형식과 도구를 제공한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.