[Paper Review] Model-Free Adaptive Optimal Control of Sequential Manufacturing Processes using Reinforcement Learning.
This paper proposes a model-free adaptive optimal control framework for sequential manufacturing processes using Q-learning-based reinforcement learning, eliminating the need for prior process modeling. By learning optimal control policies directly from real-time product quality feedback, the method achieves superior performance over model-based approaches like Model Predictive Control and Approximate Dynamic Programming in FEM-simulated deep drawing processes.
A self-learning optimal control algorithm for sequential manufacturing processes with time-discrete control actions is proposed and evaluated with simulated deep drawing processes. The necessary control model is built during consecutive process executions under optimal control via Reinforcement Learning, using the measured product quality as reward after each process execution. Prior model formation, which is required by state-of-the-art algorithms like Model Predictive Control and Approximate Dynamic Programming, is therefore obsolete. This avoids the difficulties in system identification and accurate modelling, which arise with processes subject to non-linear dynamics and stochastic influences. Also runtime complexity problems of these approaches are avoided, which arise when more complex models and larger control prediction horizons are employed. Instead of using pre-created process- and observation-models, Reinforcement Learning algorithms build functions of expected future reward during processing, which are then used for optimal process control decisions. The learning of such expectation functions is realized online by interacting with the process. The proposed algorithm also takes stochastic variations of the process conditions into consideration and is able to cope with partial observability. A method for the adaptive optimal control of partially observable fixed-horizon manufacturing processes, based on Q-learning is developed and studied. The resulting algorithm is instantiated and then evaluated by application to a time-stochastic optimal control problem in metal sheet deep drawing, where the experiments use FEM-simulated processes. The Reinforcement Learning based control shows superior results over the model-based Model Predictive Control and Approximate Dynamic Programming approaches.
Motivation & Objective
- To eliminate the need for prior process modeling in optimal control of sequential manufacturing processes.
- To address challenges in system identification and model accuracy for non-linear, stochastic processes.
- To reduce runtime complexity associated with complex models and large prediction horizons in traditional optimal control methods.
- To enable adaptive control in partially observable, time-stochastic manufacturing environments.
- To develop and evaluate a Q-learning-based algorithm for fixed-horizon optimal control in deep drawing processes.
Proposed method
- Reinforcement learning is used to learn value functions representing expected future rewards from measured product quality after each process cycle.
- The algorithm learns optimal control policies online through interaction with the physical process, without requiring pre-defined process or observation models.
- A Q-learning-based approach is employed to handle partial observability and stochastic variations in process conditions.
- The method operates in a time-discrete control framework, updating policies based on real-time feedback from simulated deep drawing processes.
- The learning process builds expectation functions for future rewards, which guide optimal control decisions without explicit system dynamics modeling.
- The algorithm is instantiated and evaluated using finite element method (FEM)-simulated deep drawing processes with stochastic inputs.
Experimental results
Research questions
- RQ1Can a model-free reinforcement learning approach achieve optimal control in sequential manufacturing processes without prior system modeling?
- RQ2How does the performance of the proposed RL-based method compare to model-based approaches like Model Predictive Control and Approximate Dynamic Programming?
- RQ3To what extent can the RL method handle stochastic variations and partial observability in manufacturing processes?
- RQ4What is the impact of eliminating system identification and model complexity on control runtime and adaptability?
- RQ5Can the RL algorithm learn effective control policies directly from product quality feedback in a simulated deep drawing process?
Key findings
- The proposed reinforcement learning-based control method outperforms both Model Predictive Control and Approximate Dynamic Programming in terms of control performance on FEM-simulated deep drawing processes.
- The method successfully learns optimal control policies without requiring prior process modeling, avoiding challenges in system identification and model accuracy.
- The algorithm effectively handles stochastic variations and partial observability in the manufacturing process, maintaining robust control performance.
- Runtime complexity is significantly reduced compared to model-based methods that rely on complex models and large prediction horizons.
- The learning process converges to optimal control policies through direct interaction with the process, using product quality as the sole reward signal.
- The results demonstrate the feasibility and superiority of model-free RL for adaptive optimal control in complex, non-linear manufacturing environments.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.