Skip to main content
QUICK REVIEW

[論文レビュー] A deep reinforcement learning model for predictive maintenance planning of road assets: Integrating LCA and LCCA

Moein Latifi, Fateme Golivand Darvishvand|arXiv (Cornell University)|Dec 20, 2021
Infrastructure Maintenance and Monitoring被引用数 9
ひとこと要約

本稿では、道路資産の予測保全計画を最適化するため、ライフサイクルアセスメント(LCA)とライフサイクルコスト分析(LCCA)を統合した深層強化学習(DRL)フレームワークを提案する。LTPPデータベースを用い、プロキシマルポリシー最適化(PPO)を用いてDRLエージェントを訓練し、経済的・環境的影響を最小限に抑えつつ、20年間の計画期間にわたり道路の状態を閾値以上に維持する最適な保全タイミングと種別を決定する。テキサス州の高速道路を対象とした事例研究において検証された。

ABSTRACT

Road maintenance planning is an integral part of road asset management. One of the main challenges in Maintenance and Rehabilitation (M&R) practices is to determine maintenance type and timing. This research proposes a framework using Reinforcement Learning (RL) based on the Long Term Pavement Performance (LTPP) database to determine the type and timing of M&R practices. A predictive DNN model is first developed in the proposed algorithm, which serves as the Environment for the RL algorithm. For the Policy estimation of the RL model, both DQN and PPO models are developed. However, PPO has been selected in the end due to better convergence and higher sample efficiency. Indicators used in this study are International Roughness Index (IRI) and Rutting Depth (RD). Initially, we considered Cracking Metric (CM) as the third indicator, but it was then excluded due to the much fewer data compared to other indicators, which resulted in lower accuracy of the results. Furthermore, in cost-effectiveness calculation (reward), we considered both the economic and environmental impacts of M&R treatments. Costs and environmental impacts have been evaluated with paLATE 2.0 software. Our method is tested on a hypothetical case study of a six-lane highway with 23 kilometers length located in Texas, which has a warm and wet climate. The results propose a 20-year M&R plan in which road condition remains in an excellent condition range. Because the early state of the road is at a good level of service, there is no need for heavy maintenance practices in the first years. Later, after heavy M&R actions, there are several 1-2 years of no need for treatments. All of these show that the proposed plan has a logical result. Decision-makers and transportation agencies can use this scheme to conduct better maintenance practices that can prevent budget waste and, at the same time, minimize the environmental impacts.

研究の動機と目的

  • 道路資産の経済的および環境的パフォーマンスを最適化する予測保全計画フレームワークの開発。
  • 道路ネットワークにおける最適な保全タイミングと処置種別の特定という課題への対処。
  • 強化学習意思決定プロセスにライフサイクルコスト分析(LCCA)とライフサイクルアセスメント(LCA)を統合すること。
  • データ駆動型で予測可能な保全スケジューリングにより、予算の無駄と環境影響を削減すること。
  • 温暖で湿潤な気候条件下における実際の高速道路事例研究でのモデルの妥当性検証。

提案手法

  • 長期間パビメントパフォーマンス(LTPP)データベース上でトレーニングされたディープニューラルネットワーク(DNN)が、強化学習の環境モデルとして機能する。
  • 収束性とサンプル効率の観点でDQNより優れるため、ポリシー学習にプロキシマルポリシー最適化(PPO)を採用する。
  • 状態空間には、国際的粗さ指数(IRI)とラッティング深さ(RD)の2つの性能指標が含まれる。データ不足のため、クラッキングメトリック(CM)は除外された。
  • 報酬関数には、paLATE 2.0を介した経済的コストとLCAを介した環境影響が統合され、ライフサイクル全体のパフォーマンスを反映する。
  • フレームワークは、サービス品質、コスト、持続可能性のバランスを考慮しながら、時間経過に伴う最適な行動をシミュレートすることで20年間の保全計画を生成する。
  • 実際の性能トレンドを用いて、温暖で湿潤な気候条件下の23 km、6車線の高速道路でモデルを検証した。

実験結果

リサーチクエスチョン

  • RQ1強化学習を用いて、道路資産の最適な保全タイミングと種別をどのように特定できるか?
  • RQ2LCAとLCCAを統合することで、道路保全意思決定の経済的および環境的パフォーマンスにどのような影響を与えるか?
  • RQ3DQNとPPOの異なる強化学習アルゴリズムは、道路保全計画において収束性とサンプル効率の観点でどのように比較されるか?
  • RQ4DRLベースのシステムは、長期的なコストと環境影響を最小限に抑えつつ、どの程度まで道路状態を閾値以上に維持できるか?
  • RQ5特定の性能指標(例:CM)におけるデータ不足は、予測保全モデルの信頼性と正確性にどのように影響を与えるか?

主な発見

  • PPOベースのDRLエージェントは収束速度とサンプル効率の点でDQNを上回り、実運用に適していることが示された。
  • モデルは、道路状態が「優良」なサービス範囲内に保たれ、劣化が最小限に抑えられた20年間の保全計画を生成した。
  • 初期段階では、道路の初期状態が良好であったため、重い保全措置は必要とされず、モデルの予測能力が裏付けられた。
  • 大規模な保全措置後、モデルは1〜2年間の間隔で治療を実施しないよう提案しており、過剰保全を回避する戦略的計画が示された。
  • 報酬関数にLCCAとLCAを統合したことで、コストと環境影響の両方のバランスの取れた最適化が達成され、時間経過とともに両方が削減された。
  • クラッキングメトリック(CM)の除外により、モデルの精度が向上した。これは、DRL応用においてデータ品質の重要性を示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。