[論文レビュー] An Opinion Mining of Text in COVID-19 Issues along with Comparative Study in ML, BERT & RNN
本稿では、COVID-19関連のバングラ語テキスト向けに、多言語感情分析システムを提案し、従来の機械学習(ML)、再帰的ニューラルネットワーク(RNN)、およびBERTベースのモデルを比較している。MLモデルでは91%の精度を達成し、ディープラーニングモデルでは79%の精度を示しており、公衆衛生危機時における低リソース言語としてのバングラ語向けに感情分析システムを展開する可能性を示している。
The global world is crossing a pandemic situation where this is a catastrophic outbreak of Respiratory Syndrome recognized as COVID-19. This is a global threat all over the 212 countries that people every day meet with mighty situations. On the contrary, thousands of infected people live rich in mountains. Mental health is also affected by this worldwide coronavirus situation. Due to this situation online sources made a communicative place that common people shares their opinion in any agenda. Such as affected news related positive and negative, financial issues, country and family crisis, lack of import and export earning system etc. different kinds of circumstances are recent trendy news in anywhere. Thus, vast amounts of text are produced within moments therefore, in subcontinent areas the same as situation in other countries and peoples opinion of text and situation also same but the language is different. This article has proposed some specific inputs along with Bangla text comments from individual sources which can assure the goal of illustration that machine learning outcome capable of building an assistive system. Opinion mining assistive system can be impactful in all language preferences possible. To the best of our knowledge, the article predicted the Bangla input text on COVID-19 issues proposed ML algorithms and deep learning models analysis also check the future reachability with a comparative analysis. Comparative analysis states a report on text prediction accuracy is 91% along with ML algorithms and 79% along with Deep Learning Models.
研究の動機と目的
- COVID-19関連のバングラ語テキストを処理できる感情分析システムの開発を目的とする。
- 従来の機械学習、RNN、BERTベースのモデルがバングラ語テキスト分類に与える性能を評価すること。
- 世界的な健康危機時における、バングラ語のような低リソース言語向けに感情分析システムを展開可能かどうかを検討すること。
- 多言語的・低リソース環境下における感情分類のモデル精度と信頼性の比較分析を提供すること。
提案手法
- COVID-19関連のオンラインソースから収集したバングラ語のコメントを収集・前処理した。
- 前処理済みバングラ語テキストに対する感情分類に、従来の機械学習モデル(例:SVM、ナイーブベイズ)を適用した。
- 同じバングラ語データセットを用いて、シーケンスモデリングと感情予測のためのRNNベースのモデルを実装した。
- 微調整を用いて事前学習済みのBERTモデル(mBERT)をバングラ語テキストの感情分類に適用し、トランスファーラーニングを活用した。
- バングラ語固有のトークン化、ストップワード除去、語形還元などの標準的なNLP前処理手順を実施した。
- すべてのモデルを標準的な分類指標を用いて評価し、主な性能指標として精度を用いた。
実験結果
リサーチクエスチョン
- RQ1従来の機械学習モデルは、COVID-19関連のバングラ語テキストの感情分類においてどの程度効果的か?
- RQ2RNNベースのモデルは、低リソースのバングラ語テキストにおける感情分類でどの程度の性能を示すか?
- RQ3この特定の低リソース環境下で、BERTベースの微調整はMLおよびRNNモデルと比較して感情予測にどの程度優れているか?
- RQ4世界的なパンデミック時における、低リソース非英語テキストを処理するさまざまなNLPモデルの相対的精度はいかがなものか?
主な発見
- 機械学習モデルは、COVID-19関連のバングラ語テキストの感情分類において、最高の91%の精度を達成した。
- RNNベースのモデルは、同じタスクで79%の感情分類精度を示し、中程度の性能であった。
- BERTベースのモデルは高い性能を示したが、この特定の低リソース環境下では従来のMLモデルに劣っていた。
- 比較分析により、学習データが限られる状況下では、従来のMLモデルがディープラーニングモデルよりも感情分析においてより効果的であることが確認された。
- 本研究は、トランスファーラーニングと微調整を活用することで、バングラ語のような低リソース言語向けに感情分析システムを効果的に適合可能であることを示している。
- 結果から、公衆衛生緊急時における非英語・低リソース文脈での感情分析において、MLモデルが依然として実用的で正確な選択肢であると考えられる。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。