Skip to main content
QUICK REVIEW

[论文解读] Condition-Based Production for Stochastically Deteriorating Systems: Optimal Policies and Learning

Collin Drent, Melvin Drent|arXiv (Cornell University)|Aug 15, 2023
Reliability and Maintenance OptimizationEngineering被引用 3
一句话总结

本文提出了一种用于随机退化系统的基于状态的生产和维护的联合优化框架,采用连续时间马尔可夫决策过程和贝叶斯学习,根据实时状态监控动态调整生产速率。结果表明,将运行生产控制与战术维护规划相结合,相比顺序方法,利润最高可提升21%;且贝叶斯学习策略的表现几乎与已知真实参数的Oracle策略相当,尤其在bang-bang控制制度下表现更优。

ABSTRACT

Production systems deteriorate stochastically due to usage and may eventually break down, resulting in high maintenance costs at scheduled maintenance moments. This deterioration behavior is affected by the system's production rate. While producing at a higher rate generates more revenue, the system may also deteriorate faster. Production should thus be controlled dynamically to trade-off deterioration and revenue accumulation in between maintenance moments. We study systems for which the relation between production and deterioration is known and the same for each system as well as systems for which this relation differs from system to system and needs to be learned on-the-fly. The decision problem is to find the optimal production policy given planned maintenance moments (operational) and the optimal interval length between such maintenance moments (tactical). For systems with a known production-deterioration relation, we cast the operational decision problem as a continuous-time Markov decision process and prove that the optimal policy has intuitive monotonic properties. We also present sufficient conditions for the optimality of bang-bang policies and we partially characterize the structure of the optimal interval length, thereby enabling efficient joint optimization of the operational and tactical decision problem. For systems that exhibit variability in their production-deterioration relations, we propose a Bayesian procedure to learn the unknown deterioration rate under any production policy. Our extensive numerical study indicates significant profit increases of our approaches compared to the state-of-the-art.

研究动机与目标

  • 解决在使用相关退化特性的生产系统中,动态平衡收益生成与系统退化的问题。
  • 在已知生产-退化关系的前提下,利用连续时间马尔可夫决策过程(CT-MDPs)开发最优运行生产策略。
  • 制定并求解战术决策问题,即确定与基于状态的生产相集成的最优维护间隔。
  • 将框架扩展至具有单元间退化行为差异的系统,其中生产-退化关系未知,需实时学习。
  • 提出并评估一种贝叶斯学习程序,以在任意生产策略下实时估计未知的退化速率。

提出的方法

  • 将运行生产控制问题建模为连续时间马尔可夫决策过程(CT-MDP),状态空间由系统状态和下次维护的剩余时间定义。
  • 证明最优策略在状态和时间上表现出直观的单调性特征,并推导出bang-bang控制(即生产速率为0或最大值)的充分条件。
  • 开发一种贝叶斯学习程序,实时估计每个系统未知的基础退化速率λ,基于任意生产策略下的观测状态信号更新信念。
  • 利用学习到的后验分布,在生产-退化关系未知或可变的系统中指导生产决策。
  • 联合优化维护间隔长度与运行策略,利用最优策略的结构特性实现高效计算。
  • 实现一种启发式策略(CE策略),其可调学习时域N^opt用于在探索与利用之间取得平衡,以学习退化速率。
Figure 1 : Illustration of monotonic behavior of the optimal condition-based production policy with base rate $\lambda=1$ (left) and base rate $\lambda=4$ (right). The color-bar indicates the optimal production rate.
Figure 1 : Illustration of monotonic behavior of the optimal condition-based production policy with base rate $\lambda=1$ (left) and base rate $\lambda=4$ (right). The color-bar indicates the optimal production rate.

实验结果

研究问题

  • RQ1在已知生产-退化关系的系统中,最优基于状态的生产策略具有何种结构性特征?
  • RQ2在随机退化系统中,何种条件下bang-bang生产策略为最优?
  • RQ3在整合基于状态的生产决策时,如何确定计划维护事件之间的最优间隔长度?
  • RQ4对于具有单元间差异的系统,如何实时学习未知的生产-退化关系?
  • RQ5所提出的贝叶斯学习策略与已知真实退化速率的Oracle策略相比,性能如何?

主要发现

  • 在广泛系统设置下,基于状态的生产相比静态生产策略,平均利润提升50%。
  • 将基于状态的生产与战术维护规划相结合,相比最先进的顺序方法,利润提升21%。
  • 所提出的贝叶斯学习策略表现几乎与已知真实退化速率的Oracle策略相当,尤其在bang-bang制度下表现更优。
  • 在bang-bang制度下,贝叶斯策略对退化速率的单元间差异具有鲁棒性,因其继承了与Oracle相同的策略结构。
  • 当最优策略非bang-bang时,具有N^opt > 0的贝叶斯学习策略能有效学习退化速率并保持强劲性能。
  • 当与最优运行策略结合时,最优维护间隔长度可部分表征并高效计算,从而实现联合优化。
Figure 2 : Sample paths of the optimal condition-based production policy (bottom) and controlled deterioration process (top), when the base rate $\lambda$ is 1.
Figure 2 : Sample paths of the optimal condition-based production policy (bottom) and controlled deterioration process (top), when the base rate $\lambda$ is 1.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。