Skip to main content
QUICK REVIEW

[Paper Review] Deep reinforcement learning for the control of conjugate heat transfer with application to workpiece cooling

Elie Hachem, Hassan Ghraieb|arXiv (Cornell University)|Nov 30, 2020
Heat Transfer Mechanisms65 references4 citations
TL;DR

This paper proposes a novel deep reinforcement learning (DRL) framework using a degenerate proximal policy optimization (PPO) algorithm to optimize conjugate heat transfer in fluid-structure systems, particularly for workpiece cooling. It demonstrates effective control of natural and forced convection in 2D and 3D setups, achieving improved temperature homogeneity and discovering non-intuitive optimal configurations, such as off-axis workpiece positioning under symmetric actuation.

ABSTRACT

This research gauges the ability of deep reinforcement learning (DRL) techniques to assist the control of conjugate heat transfer systems governed by the coupled Navier--Stokes and heat equations. It uses a novel, "degenerate" version of the proximal policy optimization (PPO) algorithm, intended for situations where the optimal policy to be learnt by a neural network does not depend on state, as is notably the case in optimization and open-loop control problems. The numerical reward fed to the neural network is computed with an in-house stabilized finite elements environment combining variational multi-scale (VMS) modeling of the governing equations, immerse volume method, and multi-component anisotropic mesh adaptation. Several test cases of natural and forced convection in two and three dimensions are used as testbed for developing the methodology. The approach successfully alleviates the natural convection induced enhancement of heat transfer in a two-dimensional, differentially heated square cavity controlled by piece-wise constant fluctuations of the sidewall temperature. It also proves capable of improving the homogeneity of temperature across the surface of two and three-dimensional hot workpieces under impingement cooling. Various cases are tackled, in which the position of multiple cold air injectors is optimized relative to a fixed workpiece position. The flexibility of the numerical framework makes it tractable to solve also the inverse problem, i.e., to optimize the workpiece position relative to a fixed injector distribution. The obtained results showcase the potential of the method for black-box optimization of practically meaningful computational fluid dynamics (CFD) conjugate heat transfer systems.

Motivation & Objective

  • To develop a DRL-based control strategy for conjugate heat transfer governed by coupled Navier–Stokes and heat equations.
  • To address the challenge of optimizing large, high-dimensional parameter spaces in thermal control problems with minimal prior knowledge.
  • To explore the potential of DRL in discovering unanticipated, high-performance control configurations beyond conventional design intuition.
  • To demonstrate the method’s flexibility in solving both forward and inverse control problems in thermal management.
  • To validate the approach on realistic 2D and 3D conjugate heat transfer scenarios, including impingement cooling and differentially heated cavities.

Proposed method

  • Adapts a 'degenerate' version of proximal policy optimization (PPO) where the policy network is trained without state dependence, suitable for open-loop and optimization problems.
  • Employs a custom stabilized finite element solver with variational multi-scale (VMS) modeling, immersed boundary method, and anisotropic mesh adaptation for accurate CFD simulation.
  • Uses a numerically computed reward signal derived from temperature homogeneity and thermal gradients to guide policy learning.
  • Integrates a multi-component anisotropic mesh adaptation strategy to improve solution accuracy and reduce computational cost in complex geometries.
  • Enables both forward control (optimizing injector positions) and inverse control (optimizing workpiece position) within the same framework.
  • Leverages deep neural networks to learn control policies directly from simulation data, treating the system as a black-box optimization problem.

Experimental results

Research questions

  • RQ1Can a degenerate PPO algorithm effectively learn optimal control policies in conjugate heat transfer systems without state-dependent action selection?
  • RQ2To what extent can DRL improve temperature uniformity across 2D and 3D hot workpieces under impingement cooling?
  • RQ3Does DRL reveal non-intuitive or counterintuitive optimal configurations, such as asymmetric workpiece positioning under symmetric actuation?
  • RQ4How scalable and robust is the DRL framework in handling high-dimensional parameter spaces typical of industrial thermal control problems?
  • RQ5Can the same framework efficiently solve both forward and inverse control problems in conjugate heat transfer?

Key findings

  • The degenerate PPO algorithm successfully mitigates natural convection-induced heat transfer enhancement in a 2D differentially heated cavity by optimizing sidewall temperature fluctuations.
  • The method significantly improves temperature homogeneity across the surface of 2D and 3D hot workpieces under impingement cooling by optimizing the spatial distribution of cold air injectors.
  • In the inverse problem setting, the DRL framework identifies that optimal workpiece positioning under symmetric actuation is off-center, a non-intuitive result not anticipated by symmetry-based design.
  • The approach demonstrates robustness and efficiency in high-dimensional control spaces, outperforming simple parametric sweeps in convergence and solution quality.
  • The framework is flexible enough to handle both forward and inverse control problems within the same numerical environment, enabling broader design exploration.
  • The results suggest that DRL can uncover novel, high-performance control strategies that are not accessible through traditional optimization or heuristic design approaches.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.