[论文解读] Model-Free Adaptive Optimal Control of Sequential Manufacturing Processes using Reinforcement Learning.
该论文提出了一种基于Q-learning强化学习的无模型自适应最优控制框架,用于顺序制造过程,无需事先进行过程建模。通过直接从实时产品品质反馈中学习最优控制策略,该方法在有限元法(FEM)模拟的深拉伸成形过程中,性能优于基于模型的方法(如模型预测控制和近似动态规划)。
A self-learning optimal control algorithm for sequential manufacturing processes with time-discrete control actions is proposed and evaluated with simulated deep drawing processes. The necessary control model is built during consecutive process executions under optimal control via Reinforcement Learning, using the measured product quality as reward after each process execution. Prior model formation, which is required by state-of-the-art algorithms like Model Predictive Control and Approximate Dynamic Programming, is therefore obsolete. This avoids the difficulties in system identification and accurate modelling, which arise with processes subject to non-linear dynamics and stochastic influences. Also runtime complexity problems of these approaches are avoided, which arise when more complex models and larger control prediction horizons are employed. Instead of using pre-created process- and observation-models, Reinforcement Learning algorithms build functions of expected future reward during processing, which are then used for optimal process control decisions. The learning of such expectation functions is realized online by interacting with the process. The proposed algorithm also takes stochastic variations of the process conditions into consideration and is able to cope with partial observability. A method for the adaptive optimal control of partially observable fixed-horizon manufacturing processes, based on Q-learning is developed and studied. The resulting algorithm is instantiated and then evaluated by application to a time-stochastic optimal control problem in metal sheet deep drawing, where the experiments use FEM-simulated processes. The Reinforcement Learning based control shows superior results over the model-based Model Predictive Control and Approximate Dynamic Programming approaches.
研究动机与目标
- 消除顺序制造过程中最优控制对先验过程建模的需求。
- 解决非线性、随机过程在系统辨识和模型精度方面的挑战。
- 降低传统最优控制方法中因复杂模型和大预测时域带来的运行时复杂度。
- 实现在部分可观测、时间随机的制造环境中的自适应控制。
- 开发并评估一种基于Q-learning的算法,用于深拉伸成形过程中的固定时域最优控制。
提出的方法
- 使用强化学习学习价值函数,表示每次工艺周期后从测量到的产品质量中获得的未来奖励期望值。
- 该算法通过与物理过程的在线交互学习最优控制策略,无需预定义的过程或观测模型。
- 采用基于Q-learning的方法,以处理工艺条件中的部分可观测性和随机波动。
- 该方法在时间离散控制框架下运行,基于来自模拟深拉伸成形过程的实时反馈更新控制策略。
- 学习过程构建未来奖励的期望函数,从而在无需显式系统动力学建模的情况下指导最优控制决策。
- 该算法通过使用有限元法(FEM)模拟的深拉伸成形过程(含随机输入)进行实例化和评估。
实验结果
研究问题
- RQ1无模型强化学习方法是否能在无需先验系统建模的情况下实现在顺序制造过程中的最优控制?
- RQ2所提出的基于强化学习的方法与基于模型的方法(如模型预测控制和近似动态规划)相比,性能如何?
- RQ3该强化学习方法在多大程度上能够处理制造过程中的随机波动和部分可观测性?
- RQ4消除系统辨识和模型复杂度对控制运行时间和适应性有何影响?
- RQ5强化学习算法是否能直接从模拟深拉伸成形过程的产品质量反馈中学习到有效的控制策略?
主要发现
- 所提出的基于强化学习的控制方法在有限元法(FEM)模拟的深拉伸成形过程中,其控制性能优于模型预测控制(MPC)和近似动态规划(ADP)。
- 该方法成功在无需先验过程建模的情况下学习到最优控制策略,避免了系统辨识和模型精度方面的挑战。
- 该算法有效处理了制造过程中的随机波动和部分可观测性,保持了鲁棒的控制性能。
- 与依赖复杂模型和大预测时域的传统基于模型方法相比,运行时复杂度显著降低。
- 学习过程通过与工艺过程的直接交互实现收敛,仅使用产品品质作为唯一奖励信号。
- 结果表明,无模型强化学习在复杂、非线性制造环境中的自适应最优控制方面具有可行性与优越性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。