Skip to main content
QUICK REVIEW

[論文レビュー] Condition-Based Production for Stochastically Deteriorating Systems: Optimal Policies and Learning

Collin Drent, Melvin Drent|arXiv (Cornell University)|Aug 15, 2023
Reliability and Maintenance OptimizationEngineering被引用数 3
ひとこと要約

本稿では、連続時間マルコフ決定過程とベイズ学習を用いて、リアルタイムの状態監視に基づき動的に生産レートを調整する、確率的劣化を示すシステムにおける状態依存生産と保守の共同最適化フレームワークを提案する。実験では、運用的生産制御と戦略的保守計画の統合が、逐次的手法と比較して利益を最大21%向上させることを示しており、ベイズ学習ポリシーは特にバン・バン制御制度下でオラクルポリシーとほぼ同等の性能を発揮することがわかった。

ABSTRACT

Production systems deteriorate stochastically due to usage and may eventually break down, resulting in high maintenance costs at scheduled maintenance moments. This deterioration behavior is affected by the system's production rate. While producing at a higher rate generates more revenue, the system may also deteriorate faster. Production should thus be controlled dynamically to trade-off deterioration and revenue accumulation in between maintenance moments. We study systems for which the relation between production and deterioration is known and the same for each system as well as systems for which this relation differs from system to system and needs to be learned on-the-fly. The decision problem is to find the optimal production policy given planned maintenance moments (operational) and the optimal interval length between such maintenance moments (tactical). For systems with a known production-deterioration relation, we cast the operational decision problem as a continuous-time Markov decision process and prove that the optimal policy has intuitive monotonic properties. We also present sufficient conditions for the optimality of bang-bang policies and we partially characterize the structure of the optimal interval length, thereby enabling efficient joint optimization of the operational and tactical decision problem. For systems that exhibit variability in their production-deterioration relations, we propose a Bayesian procedure to learn the unknown deterioration rate under any production policy. Our extensive numerical study indicates significant profit increases of our approaches compared to the state-of-the-art.

研究の動機と目的

  • 使用に伴う劣化を示す生産システムにおいて、収益の創出とシステム劣化の動的バランスをとる課題に取り組む。
  • 連続時間マルコフ決定過程(CT-MDP)を用いて、生産-劣化関係が既知の下で最適な運用的生産ポリシーを構築する。
  • 状態依存生産と統合された最適な保守インターバルを特定する戦略的意思決定問題を定式化・解法する。
  • 各ユニットの劣化行動にばらつきが生じるシステムに拡張し、生産-劣化関係が未知であり、リアルタイムで学習する必要がある状況を扱う。
  • 任意の生産ポリシー下で、未知の劣化率をリアルタイムで推定するためのベイズ学習手順を提案・評価する。

提案手法

  • 運用的生産制御問題を、システムの状態と次回保守までの時間で定義される状態空間を有する連続時間マルコフ決定過程(CT-MDP)としてモデル化する。
  • 最適ポリシーが状態および時間に関して直感的な単調性を示すことを証明し、バン・バン制御(すなわち、0または最大レートでの生産)が最適となる十分条件を導出する。
  • 任意の生産ポリシー下で観測された状態信号に基づき、信念を更新することで、各システムの未知の基本劣化率λをリアルタイムで推定するベイズ学習手順を開発する。
  • 学習された後方分布を用いて、未知または変動する生産-劣化関係を有するシステムにおける生産意思決定を支援する。
  • 最適運用ポリシーと連携して保守インターバル長を最適化し、最適ポリシーの構造的性質を活用して計算を効率化する。
  • 探索と活用のバランスを取るために、調整可能な学習期間長N^optを持つヒューリスティックポリシー(CEポリシー)を実装する。
Figure 1 : Illustration of monotonic behavior of the optimal condition-based production policy with base rate $\lambda=1$ (left) and base rate $\lambda=4$ (right). The color-bar indicates the optimal production rate.
Figure 1 : Illustration of monotonic behavior of the optimal condition-based production policy with base rate $\lambda=1$ (left) and base rate $\lambda=4$ (right). The color-bar indicates the optimal production rate.

実験結果

リサーチクエスチョン

  • RQ1生産-劣化関係が既知のシステムにおいて、最適な状態依存生産ポリシーの構造的性質は何か?
  • RQ2確率的劣化を示すシステムにおいて、バン・バン生産ポリシーが最適となる条件は何か?
  • RQ3状態依存生産意思決定を統合する場合、計画的保守イベント間の最適なインターバル長はどのように決定できるか?
  • RQ4ユニットごとのばらつきを示すシステムにおいて、未知の生産-劣化関係をリアルタイムでどのように学習できるか?
  • RQ5提案されたベイズ学習ポリシーの性能は、真の劣化率を既知とするオラクルポリシーと比べてどの程度か?

主な発見

  • 幅広いシステム設定において、状態依存生産は静的生産ポリシーと比較して平均利益を50%向上させる。
  • 状態依存生産と戦略的保守計画の統合により、最先端の逐次的手法と比較して利益が21%向上する。
  • 提案されたベイズ学習ポリシーは、特にバン・バン制御制度下でオラクルポリシーとほぼ同等の性能を発揮する。
  • バン・バン制御制度下では、劣化率のユニット間ばらつきに対してもベイズポリシーの性能が安定しており、オラクルと同一のポリシー構造を継承している。
  • 最適ポリシーがバン・バンでない場合、N^opt > 0 のベイズ学習ポリシーは劣化率を効果的に学習し、高い性能を維持する。
  • 最適運用ポリシーと連携した場合、最適保守インターバル長は部分的に特徴づけられ、効率的に計算可能となり、共同最適化が可能になる。
Figure 2 : Sample paths of the optimal condition-based production policy (bottom) and controlled deterioration process (top), when the base rate $\lambda$ is 1.
Figure 2 : Sample paths of the optimal condition-based production policy (bottom) and controlled deterioration process (top), when the base rate $\lambda$ is 1.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。