Skip to main content
QUICK REVIEW

[论文解读] Application of twin delayed deep deterministic policy gradient learning for the control of transesterification process.

Tanuja Joshi, Shikhar Makker|arXiv (Cornell University)|Feb 25, 2021
Advanced Control Systems Optimization参考文献 35被引用 7
一句话总结

本研究将双延迟深度确定性策略梯度(TD3)强化学习应用于生物柴油生产的非线性间歇酯交换过程控制。TD3智能体成功学习到连续控制策略,展现出在工业生物燃料过程中实现人工智能驱动优化的巨大潜力。

ABSTRACT

The persistent depletion of fossil fuels has encouraged mankind to look for alternatives fuels that are renewable and environment-friendly. One of the promising and renewable alternatives to fossil fuels is bio-diesel produced by means of the batch transesterification process. Control of the batch transesterification process is difficult due to its complex and non-linear dynamics. It is expected that some of these challenges can be addressed by developing control strategies that directly interact with the process and learning from the experiences. To achieve the same, this study explores the feasibility of reinforcement learning (RL) based control of the batch transesterification process. In particular, the present study exploits the application of twin delayed deep deterministic policy gradient (TD3) based RL for the continuous control of the batch transesterification process. These results showcase that TD3 based controller is able to control batch transesterification process and can be a promising direction towards the goal of artificial intelligence-based control in process industries.

研究动机与目标

  • 解决生物柴油生产中复杂非线性间歇酯交换过程控制的挑战。
  • 探究深度强化学习是否能够有效管理工业规模生物燃料过程中的连续控制。
  • 评估TD3作为鲁棒强化学习算法在化学间歇过程实时自适应控制中的可行性。
  • 开发一种端到端的基于强化学习的控制器,通过经验学习,无需依赖显式工艺模型。

提出的方法

  • 采用双延迟深度确定性策略梯度(TD3)算法,实现间歇酯交换过程中的连续控制。
  • 使用深度神经网络近似策略和Q值函数,实现在高维动作空间中的函数逼近。
  • 实现经验回放缓冲区,用于存储和采样转移数据,提升数据效率和训练稳定性。
  • 应用目标策略平滑和延迟策略更新,降低Q值估计中的过度估计偏差。
  • 通过与模拟或真实间歇酯交换环境交互训练智能体,以优化产率和工艺效率。
  • 采用连续动作输出,实时控制温度和催化剂用量等工艺变量。

实验结果

研究问题

  • RQ1基于TD3的强化学习能否有效控制间歇酯交换过程的非线性动态?
  • RQ2在该控制任务中,TD3算法相较于其他深度强化学习方法在稳定性与收敛性方面表现如何?
  • RQ3TD3智能体在缺乏工艺模型先验知识的情况下,能在多大程度上学习到最优控制策略?
  • RQ4在动态条件下,TD3控制器在工艺产率和鲁棒性方面的表现如何?

主要发现

  • 基于TD3的控制器成功学习到对间歇酯交换过程的控制,表现出稳定且一致的性能。
  • 与先前的DDPG方法相比,该算法在样本效率方面表现更优,且过度估计偏差显著降低。
  • 通过动态调节温度和催化剂用量等控制输入,控制器实现了高工艺产率。
  • 目标策略平滑和延迟策略更新的引入,显著提升了训练稳定性和收敛速度。
  • 结果证实,TD3是复杂非线性化工过程中连续控制的一种可行且高效的方法。
  • 本研究为过程工业中人工智能驱动控制奠定了基础,尤其适用于可再生能源应用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。