Skip to main content
QUICK REVIEW

[论文解读] Deep reinforcement learning for the control of conjugate heat transfer with application to workpiece cooling

Elie Hachem, Hassan Ghraieb|arXiv (Cornell University)|Nov 30, 2020
Heat Transfer Mechanisms参考文献 65被引用 4
一句话总结

本文提出了一种新颖的深度强化学习(DRL)框架,采用退化版的近端策略优化(PPO)算法,以优化流体-结构系统中的共轭传热,特别针对工件冷却问题。该方法在二维和三维设置中有效控制了自然对流与强制对流,实现了更优的温度均匀性,并发现了非直观的最优配置,例如在对称驱动条件下工件位置偏离中心轴的情况。

ABSTRACT

This research gauges the ability of deep reinforcement learning (DRL) techniques to assist the control of conjugate heat transfer systems governed by the coupled Navier--Stokes and heat equations. It uses a novel, "degenerate" version of the proximal policy optimization (PPO) algorithm, intended for situations where the optimal policy to be learnt by a neural network does not depend on state, as is notably the case in optimization and open-loop control problems. The numerical reward fed to the neural network is computed with an in-house stabilized finite elements environment combining variational multi-scale (VMS) modeling of the governing equations, immerse volume method, and multi-component anisotropic mesh adaptation. Several test cases of natural and forced convection in two and three dimensions are used as testbed for developing the methodology. The approach successfully alleviates the natural convection induced enhancement of heat transfer in a two-dimensional, differentially heated square cavity controlled by piece-wise constant fluctuations of the sidewall temperature. It also proves capable of improving the homogeneity of temperature across the surface of two and three-dimensional hot workpieces under impingement cooling. Various cases are tackled, in which the position of multiple cold air injectors is optimized relative to a fixed workpiece position. The flexibility of the numerical framework makes it tractable to solve also the inverse problem, i.e., to optimize the workpiece position relative to a fixed injector distribution. The obtained results showcase the potential of the method for black-box optimization of practically meaningful computational fluid dynamics (CFD) conjugate heat transfer systems.

研究动机与目标

  • 开发一种基于DRL的控制策略,用于由耦合的Navier–Stokes方程和热传导方程控制的共轭传热问题。
  • 解决在热控问题中,对大规模、高维参数空间进行优化,且先验知识极少的挑战。
  • 探索DRL在发现超越传统设计直觉的非预期高性能控制配置方面的潜力。
  • 展示该方法在解决热管理中正向与反向控制问题方面的灵活性。
  • 在真实的二维和三维共轭传热场景中验证该方法,包括冲击冷却和不同加热的腔体。

提出的方法

  • 采用一种‘退化’版本的近端策略优化(PPO),其中策略网络在无状态依赖的情况下进行训练,适用于开环控制与优化问题。
  • 使用自定义的稳定有限元求解器,结合变分多尺度(VMS)建模、浸入边界法以及各向异性网格自适应技术,以实现精确的计算流体动力学(CFD)模拟。
  • 利用从温度均匀性与热梯度中数值计算得到的奖励信号,指导策略学习。
  • 集成多分量各向异性网格自适应策略,以提高复杂几何结构中解的精度并降低计算成本。
  • 在同一框架内实现正向控制(优化喷射器位置)与反向控制(优化工件位置)。
  • 利用深度神经网络直接从仿真数据中学习控制策略,将系统视为黑箱优化问题。

实验结果

研究问题

  • RQ1退化版PPO算法是否能在无状态依赖动作选择的情况下,有效学习共轭传热系统中的最优控制策略?
  • RQ2DRL在冲击冷却条件下,能在多大程度上改善二维与三维热工件表面的温度均匀性?
  • RQ3DRL是否揭示了非直观或反直觉的最优配置,例如在对称驱动条件下工件位置的非对称布置?
  • RQ4该DRL框架在处理工业热控问题中典型的高维参数空间时,其可扩展性与鲁棒性如何?
  • RQ5同一框架是否能高效求解共轭传热中的正向与反向控制问题?

主要发现

  • 退化版PPO算法通过优化侧壁温度波动,成功抑制了二维不同加热腔体中由自然对流引起的传热增强。
  • 该方法通过优化冷空气喷射器的空间分布,在冲击冷却条件下显著提升了二维与三维热工件表面的温度均匀性。
  • 在反向问题设置中,DRL框架发现,在对称驱动条件下最优工件位置为非中心位置,这一结果非直观,未被基于对称性的设计所预期。
  • 该方法在高维控制空间中表现出鲁棒性与高效性,其收敛性与解的质量优于简单的参数扫描方法。
  • 该框架足够灵活,可在同一数值环境中处理正向与反向控制问题,从而实现更广泛的设计探索。
  • 结果表明,DRL能够发现新颖且高性能的控制策略,这些策略通过传统优化或启发式设计方法无法获得。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。