Skip to main content
QUICK REVIEW

[论文解读] Reinforcement and Deep Reinforcement Learning-based Solutions for Machine Maintenance Planning, Scheduling Policies, and Optimization

Oluwaseyi Ogunfowora, Homayoun Najjaran|arXiv (Cornell University)|Jul 7, 2023
Reliability and Maintenance Optimization被引用 5
一句话总结

本文对强化学习(RL)和深度强化学习(DRL)在机器维护规划与优化中的应用进行了全面的文献综述。提出了一个分类体系,基于系统类型、维护策略、学习方法和优化目标对现有研究进行分类,突出显示了智能维护系统领域中的关键方法、研究发现及研究空白。

ABSTRACT

Systems and machines undergo various failure modes that result in machine health degradation, so maintenance actions are required to restore them back to a state where they can perform their expected functions. Since maintenance tasks are inevitable, maintenance planning is essential to ensure the smooth operations of the production system and other industries at large. Maintenance planning is a decision-making problem that aims at developing optimum maintenance policies and plans that help reduces maintenance costs, extend asset life, maximize their availability, and ultimately ensure workplace safety. Reinforcement learning is a data-driven decision-making algorithm that has been increasingly applied to develop dynamic maintenance plans while leveraging the continuous information from condition monitoring of the system and machine states. By leveraging the condition monitoring data of systems and machines with reinforcement learning, smart maintenance planners can be developed, which is a precursor to achieving a smart factory. This paper presents a literature review on the applications of reinforcement and deep reinforcement learning for maintenance planning and optimization problems. To capture the common ideas without losing touch with the uniqueness of each publication, taxonomies used to categorize the systems were developed, and reviewed publications were highlighted, classified, and summarized based on these taxonomies. Adopted methodologies, findings, and well-defined interpretations of the reviewed studies were summarized in graphical and tabular representations to maximize the utility of the work for both researchers and practitioners. This work also highlights the research gaps, key insights from the literature, and areas for future work.

研究动机与目标

  • 为解决工业系统中维护规划的优化挑战,以降低维护成本、提高系统可用性并确保安全性。
  • 识别并分类制造与工业系统中基于强化学习的维护调度与优化方法。
  • 提供维护系统、学习方法与优化目标的结构化分类体系,以指导未来研究。
  • 总结80余项关于强化学习与深度强化学习在维护领域应用的研究中的关键方法、发现与局限性。
  • 指出研究空白,并为基于强化学习/深度强化学习的智能维护系统未来发展提供方向。

提出的方法

  • 构建了多维分类体系,用于根据以下维度对维护系统进行分类:(1) 系统类型(单体/多体),(2) 维护策略(预防性/纠正性),(3) 学习方法(基于模型/无模型),(4) 智能体架构(单智能体/多智能体),以及(5) 优化目标。
  • 从同行评审的期刊与会议中收集并分析了80余项关于强化学习与深度强化学习在维护规划与调度中应用的研究。
  • 根据研究是否使用状态监测数据、退化模型(如维纳过程、威布尔分布、伽马分布)以及剩余使用寿命预测(RUL)等 prognostic health management(PHM)能力,对研究进行分类。
  • 将每项研究映射到核心强化学习/深度强化学习算法(如DQN、DDQN、PPO、SAC、DDPG)以及多智能体强化学习(MARL),包括与元启发式或启发式方法结合的混合方法。
  • 使用图表与表格形式总结各项研究的方法论、目标与结果,以提升清晰度与实用性。
  • 识别算法选择、问题建模与性能指标方面的重复模式,以提取关键洞见与研究空白。

实验结果

研究问题

  • RQ1在维护规划与优化中,哪些强化学习与深度强化学习算法被最常使用?
  • RQ2不同的系统配置(单体 vs. 多体)与维护策略(预防性 vs. 纠正性)如何影响强化学习/深度强化学习方法的选择?
  • RQ3基于强化学习的维护系统中,主导的优化目标是什么?它们与工业关键绩效指标(KPI)如成本、可用性与可靠性之间的对齐关系如何?
  • RQ4当前强化学习/深度强化学习在维护应用中的主要局限与研究空白是什么,特别是在泛化能力、实时部署以及与PHM系统集成方面?
  • RQ5在工业维护应用中,基于模型与无模型的强化学习方法在性能与实用性方面如何比较?

主要发现

  • 大多数研究采用无模型的深度强化学习算法,如DQN、DDQN、PPO、SAC与DDPG,表明对数据驱动、端到端训练方法的强烈偏好。
  • 多智能体强化学习在多体系统中日益普及,尤其用于维护调度与资源分配的联合优化,但单智能体方法仍占主导地位。
  • 基于模型的方法使用较少,但在已知退化过程(如维纳过程或伽马过程)的场景中展现出潜力,可实现更快的收敛速度与更高的样本效率。
  • 大量研究假设具备PHM能力(如剩余使用寿命预测)作为输入,表明与预测性维护的集成是实现高效强化学习维护规划的关键前提。
  • 最常见的优化目标为成本最小化(占45%的研究)与可用性最大化(占20%),而关注盈利能力或资源优化的研究较少。
  • 尽管兴趣持续增长,但仅有极少数研究在真实工业数据上验证其模型,凸显了仿真研究与实际工业部署之间的差距。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。