Skip to main content
QUICK REVIEW

[論文レビュー] A Novel Variable Selection Method based on Frequent Pattern Tree for Real-time Traffic Accident Risk Prediction

Lei Lin, Qian Wang|arXiv (Cornell University)|Jan 20, 2017
Traffic Prediction and Management Techniques被引用数 7
ひとこと要約

本稿では、頻出パターン(FP)ツリー・アルゴリズムと新規の相対的オブジェクト純度比(ROPR)基準を組み合わせた、リアルタイムの交通事故リスクのための予測要因を特定する新規な変数選択手法を提案する。I-64、バージニア州、2005年のデータを用いた評価において、FPツリーに基づく変数選択は予測精度を向上させ、k-NNおよびベイジアン・ネットワーク・モデルにおいてランダムフォレストに基づく変数選択を上回った。特に、61.11%の正しく予測された事故、38.16%の誤検知率を達成した。

ABSTRACT

Traffic accident data are usually noisy, contain missing values, and heterogeneous. How to select the most important variables to improve real-time traffic accident risk prediction has become a concern of many recent studies. This paper proposes a novel variable selection method based on the Frequent Pattern tree (FP tree) algorithm. First, all the frequent patterns in the traffic accident dataset are discovered. Then for each frequent pattern, a new criterion, called the Relative Object Purity Ratio (ROPR) which we proposed, is calculated. This ROPR is added to the importance score of the variables that differentiates one frequent pattern from the others. To test the proposed method, a dataset was compiled from the traffic accidents records detected by only one detector on interstate highway I-64 in Virginia in 2005. This data set was then linked to other variables such as real-time traffic information and weather conditions. Both the proposed method based on the FP tree algorithm, as well as the widely utilized, random forest method, were then used to identify the important variables or the Virginia data set. The results indicate that there are some differences between the variables deemed important by the FP tree and those selected by the random forest method. Following this, two baseline models (i.e. a nearest neighbor (k-NN) method and a Bayesian network) were developed to predict accident risk based on the variables identified by both the FP tree method and the random forest method. The results show that the models based on the variable selection using the FP tree performed better than those based on the random forest method for several versions of the k-NN and Bayesian network models.The best results were derived from a Bayesian network model using variables from FP tree. That model could predict 61.11% of accidents accurately, while having a false alarm rate of 38.16%.

研究の動機と目的

  • リアルタイムリスク予測におけるノイズ多発、不完全、多様性のある交通事故データの課題に対処すること。
  • 複雑な交通データセットから最も関連性の高い予測変数を特定する新規な変数選択手法を開発すること。
  • データ駆動型でパターンに基づくアプローチを用いて、リアルタイム交通事故リスク予測モデルの精度と信頼性を向上させること。
  • 提案手法の性能をランダムフォレストなどの既存手法と比較すること。

提案手法

  • 交通事故データセットからの頻出パターンを抽出するためにFPツリー・アルゴリズムが用いられ、事故発生時における変数の同時発現を捉える。
  • 各頻出パターンに対して、新規の相対的オブジェクト純度比(ROPR)が計算され、その構成変数の区別能を定量化する。
  • ROPR値が変数の重要度スコアに統合され、高リスクと低リスクのパターンを最も効果的に区別する変数が優先順位付けされる。
  • 選択された変数を用いて、ベースラインモデル(k-近傍法(k-NN)とベイジアン・ネットワーク)を訓練し、リスク予測を行う。
  • FPツリーで選択された変数を用いたモデルの性能を、ランダムフォレストで選択された変数を用いたモデルと比較する。
  • 本手法は、I-64、バージニア州、2005年の実世界データセットを用いて評価され、リアルタイム交通および天候データが追加済みである。

実験結果

リサーチクエスチョン

  • RQ1提案されたFPツリー基盤の変数選択手法とROPRは、交通事故リスクの主要な予測要因を特定する上でどの程度有効であるか?
  • RQ2FPツリーで選択された変数を用いた予測モデルの性能は、ランダムフォレストで選択された変数を用いたモデルと比べてどうか?
  • RQ3本手法で選択された変数を用いたモデルの予測精度と誤検知率はどの程度か?
  • RQ4従来の変数選択手法と比較して、FPツリー基盤のアプローチはリアルタイム交通事故リスクの検出をどのように改善できるか?

主な発見

  • FPツリー法で変数を選択したベイジアン・ネットワーク・モデルが、交通事故の予測精度で最高の61.11%を達成した。
  • 同じベイジアン・ネットワーク・モデルは、誤検知率が38.16%であったため、検知率と誤検知のバランスが取れていることが示された。
  • k-NNおよびベイジアン・ネットワーク・モデルの複数の設定において、FPツリーで選択された変数を用いたモデルが、ランダムフォレストで選択された変数を用いたモデルを上回った。
  • FPツリー法は、ランダムフォレストとは異なる重要な変数のセットを同定したため、リスク要因に関する補完的で独自の知見が得られた。
  • ROPR基準は、事故発生時と非事故時を効果的に区別するパターンに注目することで、変数の重要度を効果的に向上させた。
  • 本手法は、特にノイズが多く複雑な交通データセットにおいて、リアルタイム事故リスク予測の分野で優れた性能を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。