[論文レビュー] Natural Hazards Twitter Dataset
本論文は、2011年から2019年までの米国主要自然災害( tornados, hurricanes, floods, blizzards, wildfires など)に関連する約49,800件の感情ラベル付きツイートから構成される、公開可能な感情ラベル付きツイッター・データセットを紹介する。本研究では感情分類のための機械学習モデルを提案し、被災対応や必須支援要件に関する一般の世論を分析することで、自動的意見マイニングを活用した人道的支援の改善を支援するリソースを提供する。
With the development of the Internet, social media has become an important channel for posting disaster-related information. Analyzing attitudes hidden in these texts, known as sentiment analysis, is crucial for the government or relief agencies to improve disaster response efficiency, but it has not received sufficient attention. This paper aims to fill this gap by focusing on investigating attitudes towards disaster response and analyzing targeted relief supplies during disaster response. The contributions of this paper are fourfold. First, we propose several machine learning models for classifying public sentiment concerning disaster-related social media data. Second, we create a natural disaster dataset with sentiment labels, which contains nearly 50,00 Twitter data about different natural disasters in the United States (e.g., a tornado in 2011, a hurricane named Sandy in 2012, a series of floods in 2013, a hurricane named Matthew in 2016, a blizzard in 2016, a hurricane named Harvey in 2017, a hurricane named Michael in 2018, a series of wildfires in 2018, and a hurricane named Dorian in 2019). We are making our dataset available to the research community: https://github.com/Dong-UTIL/Natural-Hazards-Twitter-Dataset. It is our hope that our contribution will enable the study of sentiment analysis in disaster response. Third, we focus on extracting public attitudes and analyzing the essential needs (e.g., food, housing, transportation, and medical supplies) for the public during disaster response, instead of merely targeting on studying positive or negative attitudes of the public to natural disasters. Fourth, we conduct this research from two different dimensions for a comprehensive understanding of public opinion on disaster response, since disparate hazards caused by different types of natural disasters.
研究の動機と目的
- NLP研究における感情ラベル付き災害関連ソーシャルメディアデータの不足に対処すること。
- 災害関連ツイートの一般の世論を分類するための機械学習モデルの開発。
- 被災対応に対する一般の態度の分析および必須支援要件(例:食料、医療用品、住宅)の特定。
- 異なる災害タイプや同一災害(例:複数のハリケーン)を比較することで、一般世論の多次元的理解を図ること。
- 自動化された感情および支援要件検出を可能にすることで、人道的支援団体による的確な支援配分を支援すること。
提案手法
- TwitterScraperおよびBeautifulSoup4を用いてツイートを収集し、災害や支援要件に関連するキーワード(例:'hurricane' + 'food' や 'shelters')でフィルタリング。
- 災害発生前後1週間を含むデータ収集期間を拡張し、前後イベントにおける一般の世論を捉える。
- NLPモデルの整合性と利用可能性を確保するため、英語のツイートに限定。
- バイナリ感情ラベルを割り当て:0は肯定的(例:支援への感謝)、1は否定的(例:対応の遅れによる不満)。
- 9年間にわたり10の異なる災害イベントをカバーするデータセットを構築。5つのハリケーンと5つの他の災害タイプを含む。
- GitHub経由で公開し、Twitterの利用規約に準拠したライセンスのもと、一般研究利用を可能にした。
実験結果
リサーチクエスチョン
- RQ1機械学習モデルは、災害関連ツイートの一般の世論を効果的に分類できるか?
- RQ2異なる自然災害の過程で、最も頻繁に表明された支援要件(例:食料、医療用品)は何か?
- RQ3被災対応に対する一般の世論は、災害タイプや災害イベントによってどのように変化するか?
- RQ4ソーシャルメディアデータの感情分析は、人道的支援の的確な配分をどの程度向上できるか?
- RQ5一貫した感情ラベルと支援要件関連コンテンツを備えた統合データセットは、イベント間および災害タイプ間の分析を支援できるか?
主な発見
- 本データセットは、2011年から2019年までの米国主要自然災害(ハリケーン5件、その他の災害タイプ5件)に起因する49,816件の英語ツイートから構成され、感情ラベルが付与されている。
- 肯定的世論(例:支援への感謝)と否定的世論(例:避難の遅れによる不満)を示す感情ラベル付きの例が含まれる。
- 災害タイプごとに感情が顕著に異なり、特にハリケーンや洪水のような長期にわたる災害では否定的世論が高かった。
- 食料、住宅、輸送、医療用品といった支援要件が頻繁に言及されており、物的支援への強い一般の需要が示された。
- 本データセットにより、災害間での感情傾向の比較分析が可能となり、対応フェーズにおける一般の満足度と不満のパターンが明らかになった。
- 本データセットは https://github.com/Dong-UTIL/Natural-Hazards-Twitter-Dataset で公開されており、災害NLPおよび人道的インフォーマティクス分野における再現性のある研究を支援する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。