Skip to main content
QUICK REVIEW

[論文レビュー] A Comparative Study on Crime in Denver City Based on Machine Learning and Data Mining

Md. Aminur Rab Ratul|arXiv (Cornell University)|Jan 9, 2020
Crime Patterns and Interventions参考文献 22被引用数 13
ひとこと要約

本研究では、デンバーの犯罪データセット(478,578件の事件、2014年~2019年)を対象に、複数の機械学習およびデータマイニング手法を適用し、15の犯罪カテゴリを予測する。統計的分析、可視化、分類モデル(アンサンブル手法を含む)を用いて、アンサンブルモデル4が90%を超える精度を達成し、他のモデルと比較して精度、再現率、F1スコア、ROC分析において優れた性能を示した。これにより、警察当局および犯罪予防に向けた実用的インサイトが得られた。

ABSTRACT

To ensure the security of the general mass, crime prevention is one of the most higher priorities for any government. An accurate crime prediction model can help the government, law enforcement to prevent violence, detect the criminals in advance, allocate the government resources, and recognize problems causing crimes. To construct any future-oriented tools, examine and understand the crime patterns in the earliest possible time is essential. In this paper, I analyzed a real-world crime and accident dataset of Denver county, USA, from January 2014 to May 2019, which containing 478,578 incidents. This project aims to predict and highlights the trends of occurrence that will, in return, support the law enforcement agencies and government to discover the preventive measures from the prediction rates. At first, I apply several statistical analysis supported by several data visualization approaches. Then, I implement various classification algorithms such as Random Forest, Decision Tree, AdaBoost Classifier, Extra Tree Classifier, Linear Discriminant Analysis, K-Neighbors Classifiers, and 4 Ensemble Models to classify 15 different classes of crimes. The outcomes are captured using two popular test methods: train-test split, and k-fold cross-validation. Moreover, to evaluate the performance flawlessly, I also utilize precision, recall, F1-score, Mean Squared Error (MSE), ROC curve, and paired-T-test. Except for the AdaBoost classifier, most of the algorithms exhibit satisfactory accuracy. Random Forest, Decision Tree, Ensemble Model 1, 3, and 4 even produce me more than 90% accuracy. Among all the approaches, Ensemble Model 4 presented superior results for every evaluation basis. This study could be useful to raise the awareness of peoples regarding the occurrence locations and to assist security agencies to predict future outbreaks of violence in a specific area within a particular time.

研究の動機と目的

  • デンバーにおける犯罪パターンの正確な予測モデルを開発し、警察当局および政府のリソース配分を支援すること。
  • デンバー郡の現実の事件データを用いて、高リスクの犯罪種別および場所を特定すること。
  • 15の異なる犯罪カテゴリを分類するための多様な機械学習アルゴリズムの性能を比較すること。
  • トレイン・テスト分割およびk分割交差検証を用いた複数の指標に基づくモデルの頑健性を評価すること。
  • データ駆動型予測と時空間的傾向分析を通じて、犯罪予防に向けた実用的インサイトを提供すること。

提案手法

  • 2014年1月~2019年5月のデンバーの犯罪データにおける時系列的および空間的パターンを調査するため、統計的分析およびデータ可視化を実施。
  • 10種類の分類アルゴリズムを実装:ランダムフォレスト、決定木、AdaBoost、エクストラツリー、線形判別分析、K近傍法、および4つのアンサンブルモデル。
  • 一般化および安定性を確保するため、モデルの評価にトレイン・テスト分割およびk分割交差検証を用いた。
  • 精度、再現率、F1スコア、平均二乗誤差(MSE)、ROC曲線、および対応t検定を用いて性能を測定した。
  • 全評価指標において一貫した優位性を示したモデルを最良のモデルとして選定した。
  • 最終モデルを用いて、予測傾向を特定し、戦略的犯罪予防計画を支援した。

実験結果

リサーチクエスチョン

  • RQ1どの機械学習モデルがデンバーの15の異なる犯罪カテゴリを予測する際に最高の精度を達成するか?
  • RQ2デンバーの犯罪データセットにおいて、異なるアンサンブルおよびベース分類器は、精度、再現率、F1スコアの観点でどのように比較されるか?
  • RQ32014年から2019年までのデンバーの犯罪発生において、最も顕著な時系列的および空間的パターンは何か?
  • RQ4トレイン・テスト分割およびk分割交差検証といった異なる評価戦略において、予測の頑健性はどの程度か?
  • RQ5モデルは、予防的警察活動のための高リスクの犯罪種別および場所を信頼性高く特定できるか?

主な発見

  • アンサンブルモデル4は、全評価指標において最高の精度を達成し、他のすべてのモデルを上回った。
  • ランダムフォレスト、決定木、および他の3つのアンサンブルモデルは90%を超える精度を達成し、強力な予測性能を示した。
  • AdaBoostを除き、すべてのモデルが満足のいく性能を示し、高い精度、再現率、F1スコアを達成した。
  • トレイン・テスト分割およびk分割交差検証の両方において、モデルの性能は一貫して高く、頑健性が確認された。
  • ROC曲線分析により、特にアンサンブルモデル4において高い識別力が確認された。
  • 対応t検定により、統計的に有意な性能差が確認され、アンサンブルモデル4が他のモデルを顕著に上回った。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。