Skip to main content
QUICK REVIEW

[Paper Review] Controlling Rayleigh-Bénard convection via Reinforcement Learning

Gerben I. Beintema, Alessandro Corbetta|TU/e Research Portal|Mar 31, 2020
Model Reduction and Neural Networks64 references75 citations
TL;DR

The paper demonstrates that reinforcement learning can significantly suppress convection in a 2D Rayleigh–Bénard system by modulating bottom boundary temperatures, outperforming linear controls and increasing the controllable Ra threshold.

ABSTRACT

Thermal convection is ubiquitous in nature as well as in many industrial applications. The identification of effective control strategies to, e.g., suppress or enhance the convective heat exchange under fixed external thermal gradients is an outstanding fundamental and technological issue. In this work, we explore a novel approach, based on a state-of-the-art Reinforcement Learning (RL) algorithm, which is capable of significantly reducing the heat transport in a two-dimensional Rayleigh-Bénard system by applying small temperature fluctuations to the lower boundary of the system. By using numerical simulations, we show that our RL-based control is able to stabilize the conductive regime and bring the onset of convection up to a Rayleigh number $Ra_c \approx 3 \cdot 10^4$, whereas in the uncontrolled case it holds $Ra_{c}=1708$. Additionally, for $Ra > 3 \cdot 10^4$, our approach outperforms other state-of-the-art control algorithms reducing the heat flux by a factor of about $2.5$. In the last part of the manuscript, we address theoretical limits connected to controlling an unstable and chaotic dynamics as the one considered here. We show that controllability is hindered by observability and/or capabilities of actuating actions, which can be quantified in terms of characteristic time delays. When these delays become comparable with the Lyapunov time of the system, control becomes impossible.

Motivation & Objective

  • Motivate the control of thermally driven flows and heat transport in Rayleigh–Bénard convection.
  • Develop and compare active control strategies to suppress convection at fixed Rayleigh number.
  • Demonstrate that RL-based control outperforms linear controllers in higher-Rayleigh regimes.
  • Explore theoretical limits to controllability due to observability and actuation delays in chaotic systems.

Proposed method

  • Model the 2D Rayleigh–Bénard system with a lattice Boltzmann method in BGK form (D2Q9 velocity, D2Q4 temperature).
  • Define control by imposing bottom boundary temperature fluctuations with a bounded amplitude.
  • Compare linear PD control against an RL-based controller that outputs bottom-boundary temperature profiles.
  • Use a state space built from temperature/velocity probes over a grid to feed an MLP-based policy in a PPO RL framework.
  • Discretize actions as piece-wise constant temperature profiles on 10 segments with binary levels, normalized to satisfy constraints.
  • Evaluate performance via time-averaged Nusselt number Nu and its instantaneous form Nu_inst.

Experimental results

Research questions

  • RQ1Can RL-based control reduce convective heat transport in RBC at fixed Rayleigh number more effectively than linear control methods?
  • RQ2What are the achievable increases in the critical Rayleigh number under RL versus linear control?
  • RQ3How do control delays and observability affect the feasibility of stabilizing or suppressing RBC in chaotic regimes?
  • RQ4What flow structures emerge under RL control that lead to reduced heat transport?

Key findings

  • RL control increases the critical Rayleigh number from ~1e3 (uncontrolled) to ~1e4 (linear) and ~3e4 (RL).
  • For Ra > 3e4, RL control reduces the time-averaged Nu by about 2.5 compared to the uncontrolled case, outperforming linear methods which achieve ~1.5 reductions up to Ra < 1e6.
  • RL control stabilizes the conductive state up to Ra ~ 3e4 and, at higher Ra, achieves a stationary or reduced Nu rather than periodic flows typical of linear control.
  • RL induces flow configurations resembling a double-cell regime, effectively reducing heat transport by modifying the convection structure; this is observed up to Ra ~ 1e5, with diminishing but still present effects at Ra ~ 1e6.
  • Training times vary with Ra, from under an hour on a V100 for Ra ≲ 1e5 to ~150 hours for Ra ≳ 1e6, due to increased chaoticity.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.