[Paper Review] Condition-Based Production for Stochastically Deteriorating Systems: Optimal Policies and Learning
This paper proposes a joint optimization framework for condition-based production and maintenance in stochastically deteriorating systems, using continuous-time Markov decision processes and Bayesian learning to dynamically adjust production rates based on real-time condition monitoring. It demonstrates that integrating operational production control with tactical maintenance planning increases profits by up to 21% compared to sequential approaches, with a Bayesian learning policy performing nearly as well as an Oracle policy, especially under bang-bang control regimes.
Production systems deteriorate stochastically due to usage and may eventually break down, resulting in high maintenance costs at scheduled maintenance moments. This deterioration behavior is affected by the system's production rate. While producing at a higher rate generates more revenue, the system may also deteriorate faster. Production should thus be controlled dynamically to trade-off deterioration and revenue accumulation in between maintenance moments. We study systems for which the relation between production and deterioration is known and the same for each system as well as systems for which this relation differs from system to system and needs to be learned on-the-fly. The decision problem is to find the optimal production policy given planned maintenance moments (operational) and the optimal interval length between such maintenance moments (tactical). For systems with a known production-deterioration relation, we cast the operational decision problem as a continuous-time Markov decision process and prove that the optimal policy has intuitive monotonic properties. We also present sufficient conditions for the optimality of bang-bang policies and we partially characterize the structure of the optimal interval length, thereby enabling efficient joint optimization of the operational and tactical decision problem. For systems that exhibit variability in their production-deterioration relations, we propose a Bayesian procedure to learn the unknown deterioration rate under any production policy. Our extensive numerical study indicates significant profit increases of our approaches compared to the state-of-the-art.
Motivation & Objective
- Address the challenge of dynamically balancing revenue generation and system deterioration in production systems with usage-dependent degradation.
- Develop optimal operational production policies under known production-deterioration relationships using continuous-time Markov decision processes (CT-MDPs).
- Formulate and solve the tactical decision problem of determining optimal maintenance intervals that integrate with condition-based production.
- Extend the framework to systems with unit-to-unit variability in deterioration behavior, where the production-deterioration relationship is unknown and must be learned in real time.
- Propose and evaluate a Bayesian learning procedure to estimate unknown deterioration rates on-the-fly under any production policy.
Proposed method
- Model the operational production control problem as a continuous-time Markov decision process (CT-MDP), with state space defined by system condition and time to next maintenance.
- Prove that the optimal policy exhibits intuitive monotonicity properties in condition and time, and derive sufficient conditions for bang-bang control (i.e., production at 0 or maximum rate).
- Develop a Bayesian learning procedure to estimate the unknown base deterioration rate λ for each system in real time, updating beliefs based on observed condition signals under any production policy.
- Use the learned posterior distributions to guide production decisions in systems with unknown or variable production-deterioration relations.
- Optimize the maintenance interval length jointly with the operational policy, leveraging structural properties of the optimal policy to enable efficient computation.
- Implement a heuristic policy (CE policy) with a tunable learning horizon N^opt to balance exploration and exploitation in learning the deterioration rate.

Experimental results
Research questions
- RQ1What are the structural properties of the optimal condition-based production policy in systems with known production-deterioration relationships?
- RQ2Under what conditions is a bang-bang production policy optimal for stochastically deteriorating systems?
- RQ3How can the optimal interval length between planned maintenance events be determined when integrating condition-based production decisions?
- RQ4How can the unknown production-deterioration relationship be learned in real time for systems with unit-to-unit variability?
- RQ5How does the performance of the proposed Bayesian learning policy compare to an Oracle policy that knows the true deterioration rate?
Key findings
- Condition-based production increases average profits by 50% compared to static production policies across a wide range of system settings.
- Integrating condition-based production with tactical maintenance planning increases profits by 21% compared to the state-of-the-art sequential approach.
- The proposed Bayesian learning policy performs nearly as well as an Oracle policy that knows the true deterioration rate, especially in the bang-bang regime.
- In the bang-bang regime, the performance of the Bayesian policy is robust to unit-to-unit variability in deterioration rates, as it inherits the same policy structure as the Oracle.
- When the optimal policy is not bang-bang, the Bayesian learning policy with N^opt > 0 effectively learns the deterioration rate and maintains strong performance.
- The optimal maintenance interval length can be partially characterized and efficiently computed when combined with the optimal operational policy, enabling joint optimization.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.