[논문 리뷰] Natural Hazards Twitter Dataset
이 논문은 2011~2019년 동안 미국에서 발생한 주요 자연재해( tornados, 히러케인, 홍수, blizzards, wildfulres 포함)와 관련된 약 49,800건의 감성 레이블이 부여된 공개적인 트위터 데이터셋을 소개한다. 감성 분류를 위한 기계학습 모델을 제안하고 재해 대응 및 필수 구호 수요에 대한 대중의 감성 분석을 수행하여, 자동화된 의견 마이닝을 통해 인도적 대응을 향상시키는 데 자원을 제공한다.
With the development of the Internet, social media has become an important channel for posting disaster-related information. Analyzing attitudes hidden in these texts, known as sentiment analysis, is crucial for the government or relief agencies to improve disaster response efficiency, but it has not received sufficient attention. This paper aims to fill this gap by focusing on investigating attitudes towards disaster response and analyzing targeted relief supplies during disaster response. The contributions of this paper are fourfold. First, we propose several machine learning models for classifying public sentiment concerning disaster-related social media data. Second, we create a natural disaster dataset with sentiment labels, which contains nearly 50,00 Twitter data about different natural disasters in the United States (e.g., a tornado in 2011, a hurricane named Sandy in 2012, a series of floods in 2013, a hurricane named Matthew in 2016, a blizzard in 2016, a hurricane named Harvey in 2017, a hurricane named Michael in 2018, a series of wildfires in 2018, and a hurricane named Dorian in 2019). We are making our dataset available to the research community: https://github.com/Dong-UTIL/Natural-Hazards-Twitter-Dataset. It is our hope that our contribution will enable the study of sentiment analysis in disaster response. Third, we focus on extracting public attitudes and analyzing the essential needs (e.g., food, housing, transportation, and medical supplies) for the public during disaster response, instead of merely targeting on studying positive or negative attitudes of the public to natural disasters. Fourth, we conduct this research from two different dimensions for a comprehensive understanding of public opinion on disaster response, since disparate hazards caused by different types of natural disasters.
연구 동기 및 목표
- NLP 연구를 위한 감성 레이블이 부여된 재해 관련 소셜미디어 데이터의 부족을 해결하기 위해.
- 재해 관련 트위터 콘텐츠의 대중 감성 분류를 위한 기계학습 모델을 개발하기 위해.
- 재해 대응에 대한 대중의 태도를 분석하고 필수 구호 수요(예: 식량, 의료물자, 주거)를 특정하기 위해.
- 다양한 재해 유형과 동일한 재해(예: 여러 히러케인)를 비교함으로써 대중 의견의 다차원적 이해를 제공하기 위해.
- 자동화된 감성 및 수요 탐지 기능을 통해 인도적 구호 조직이 타겟팅된 지원을 가능하게 하기 위해.
제안 방법
- TwitterScraper와 BeautifulSoup4를 사용하여 트윗을 수집하고, 재해 및 구호 수요 관련 키워드(예: 'hurricane' + 'food' 또는 'shelters')로 필터링함.
- 재해 발생 전후로 일주일씩 데이터 수집 윈도우를 연장하여 사전 및 사후 대중의 감성을 포괄함.
- 데이터 일관성과 NLP 모델의 사용 가능성을 확보하기 위해 영어 트윗에 집중함.
- 이元 감성 레이블을 할당: 0은 긍정(예: 구호에 감사)을, 1은 부정(예: 대응 부족으로 인한 분노)을 의미함.
- 9년에 걸쳐 10개의 고유한 재해 사건을 포함한 데이터셋을 구축함. 이 중 5개는 히러케인, 5개는 기타 재해 유형임.
- GitHub를 통해 Twitter의 이용 약관을 준수하는 라이선스로 공개하여 연구 목적의 공개 접근 가능하게 제공함.
실험 결과
연구 질문
- RQ1기계학습 모델은 재해 관련 트위터 콘텐츠의 대중 감성을 효과적으로 분류할 수 있는가?
- RQ2다양한 유형의 자연재해 기간 동안 가장 자주 표현된 공중 수요(예: 식량, 의료물자)는 무엇인가?
- RQ3다양한 재해 유형과 사건 간에 재해 대응에 대한 대중 감성은 어떻게 달라지는가?
- RQ4소셜미디어 데이터의 감성 분석이 인도적 구호 지원의 타겟팅을 얼마나 향상시킬 수 있는가?
- RQ5일관된 감성 및 수요 관련 콘텐츠가 레이블이 부여된 단일 데이터셋이 사건 간 및 재해 유형 간 분석을 지원할 수 있는가?
주요 결과
- 데이터셋은 2011년에서 2019년 사이에 미국에서 발생한 10건의 주요 자연재해(5건의 히러케인 및 5건의 기타 재해 유형 포함)에서 수집된 49,816건의 영어 트윗을 포함한다.
- 데이터셋에는 긍정 감성(예: 구호에 감사)과 부정 감성(예: 대피 지연로 인한 분노)을 보여주는 감성 레이블이 부여된 예시가 포함되어 있다.
- 재해 유형 간 감성은 뚜렷하게 다름. 특히 히러케인과 홍수처럼 장기 지속되는 재해 동안 더 높은 부정 감성이 관찰됨.
- 식량, 주거, 교통, 의료물자 등의 구호 수요가 자주 언급되어 물류 지원에 대한 강한 대중 수요가 있음을 시사함.
- 데이터셋을 통해 재해 간 감성 추세의 비교 분석이 가능하여, 대응 단계에서의 대중 만족도 및 분노 패턴을 파악할 수 있음.
- 데이터셋은 https://github.com/Dong-UTIL/Natural-Hazards-Twitter-Dataset 에서 공개되어 있으며, 재해 NLP 및 인도적 인포매틱스 분야에서 재현 가능한 연구를 지원함.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.