[论文解读] A deep reinforcement learning model for predictive maintenance planning of road assets: Integrating LCA and LCCA
该论文提出了一种深度强化学习(DRL)框架,整合了生命周期评估(LCA)和生命周期成本分析(LCCA),以优化道路资产的预测性维护规划。基于LTPP数据库,采用近端策略优化(PPO)训练DRL智能体,确定最优维护时机和类型,使经济与环境影响最小化,同时在20年规划期内保持道路状况高于阈值,案例研究基于德克萨斯州的一条高速公路。
Road maintenance planning is an integral part of road asset management. One of the main challenges in Maintenance and Rehabilitation (M&R) practices is to determine maintenance type and timing. This research proposes a framework using Reinforcement Learning (RL) based on the Long Term Pavement Performance (LTPP) database to determine the type and timing of M&R practices. A predictive DNN model is first developed in the proposed algorithm, which serves as the Environment for the RL algorithm. For the Policy estimation of the RL model, both DQN and PPO models are developed. However, PPO has been selected in the end due to better convergence and higher sample efficiency. Indicators used in this study are International Roughness Index (IRI) and Rutting Depth (RD). Initially, we considered Cracking Metric (CM) as the third indicator, but it was then excluded due to the much fewer data compared to other indicators, which resulted in lower accuracy of the results. Furthermore, in cost-effectiveness calculation (reward), we considered both the economic and environmental impacts of M&R treatments. Costs and environmental impacts have been evaluated with paLATE 2.0 software. Our method is tested on a hypothetical case study of a six-lane highway with 23 kilometers length located in Texas, which has a warm and wet climate. The results propose a 20-year M&R plan in which road condition remains in an excellent condition range. Because the early state of the road is at a good level of service, there is no need for heavy maintenance practices in the first years. Later, after heavy M&R actions, there are several 1-2 years of no need for treatments. All of these show that the proposed plan has a logical result. Decision-makers and transportation agencies can use this scheme to conduct better maintenance practices that can prevent budget waste and, at the same time, minimize the environmental impacts.
研究动机与目标
- 开发一种预测性维护规划框架,以同时优化道路资产的经济与环境表现。
- 解决确定道路网络最优维护时机与处理类型的挑战。
- 将生命周期成本分析(LCCA)与生命周期评估(LCA)整合进强化学习决策过程。
- 通过数据驱动的预测性维护调度,减少预算浪费与环境影响。
- 在温湿气候条件下,于真实高速公路案例研究中验证模型。
提出的方法
- 基于长期路面性能(LTPP)数据库训练深度神经网络(DNN),作为强化学习环境模型。
- 强化学习智能体采用近端策略优化(PPO)进行策略学习,因其在收敛性与样本效率方面优于深度Q网络(DQN)。
- 状态空间包含两个性能指标:国际粗糙度指数(IRI)与车辙深度(RD),因数据不足而排除开裂指标(CM)。
- 奖励函数结合了经济成本(通过paLATE 2.0)与环境影响(通过LCA),以反映总生命周期表现。
- 通过时间序列模拟最优动作,生成20年维护计划,平衡服务质量、成本与可持续性。
- 在德克萨斯州一条23公里、六车道的高速公路(温湿气候)上验证模型,采用真实性能趋势数据。
实验结果
研究问题
- RQ1强化学习如何用于确定道路资产最优维护时机与类型?
- RQ2集成LCA与LCCA对道路维护决策的经济与环境表现有何影响?
- RQ3在道路维护规划中,不同强化学习算法(DQN与PPO)在收敛性与样本效率方面如何比较?
- RQ4基于DRL的系统在多大程度上可维持道路状况高于阈值,同时最小化长期成本与环境影响?
- RQ5某些性能指标(如CM)的数据稀缺如何影响预测性维护模型的可靠性与准确性?
主要发现
- 基于PPO的DRL智能体在收敛速度与样本效率方面优于DQN,更适合实际部署。
- 模型生成的20年维护计划中,道路状况始终保持在‘优良’服务范围内,退化程度极低。
- 在初期阶段,由于道路初始状况良好,无需实施重型维护,体现了模型的预测能力。
- 重大维护措施后,模型建议1–2年间隔内不进行处理,表明具备战略性规划能力,避免过度维护。
- 在奖励函数中整合LCCA与LCA,实现了成本与环境影响的平衡优化,二者随时间均有所降低。
- 因数据不足而排除开裂指标(CM)反而提升了模型准确性,凸显了数据质量在DRL应用中的重要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。