[论文解读] Controlling Rayleigh-Bénard convection via Reinforcement Learning
该论文表明通过调节底边界温度,强化学习能显著抑制二维雷利–本尼德系统中的对流,优于线性控制并提高可控雷数阈值。
Thermal convection is ubiquitous in nature as well as in many industrial applications. The identification of effective control strategies to, e.g., suppress or enhance the convective heat exchange under fixed external thermal gradients is an outstanding fundamental and technological issue. In this work, we explore a novel approach, based on a state-of-the-art Reinforcement Learning (RL) algorithm, which is capable of significantly reducing the heat transport in a two-dimensional Rayleigh-Bénard system by applying small temperature fluctuations to the lower boundary of the system. By using numerical simulations, we show that our RL-based control is able to stabilize the conductive regime and bring the onset of convection up to a Rayleigh number $Ra_c \approx 3 \cdot 10^4$, whereas in the uncontrolled case it holds $Ra_{c}=1708$. Additionally, for $Ra > 3 \cdot 10^4$, our approach outperforms other state-of-the-art control algorithms reducing the heat flux by a factor of about $2.5$. In the last part of the manuscript, we address theoretical limits connected to controlling an unstable and chaotic dynamics as the one considered here. We show that controllability is hindered by observability and/or capabilities of actuating actions, which can be quantified in terms of characteristic time delays. When these delays become comparable with the Lyapunov time of the system, control becomes impossible.
研究动机与目标
- Motivate the control of thermally driven flows and heat transport in Rayleigh–Bénard convection.
- Develop and compare active control strategies to suppress convection at fixed Rayleigh number.
- Demonstrate that RL-based control outperforms linear controllers in higher-Rayleigh regimes.
- Explore theoretical limits to controllability due to observability and actuation delays in chaotic systems.
提出的方法
- 用 BGK 形式的格子玻尔兹曼法对二维雷利–本尼德系统进行建模(D2Q9速度,D2Q4温度)。
- 通过对底边界温度扰动设定有限幅度来定义控制。
- 将线性 PD 控制与输出底边界温度轮廓的 RL 控制器进行比较。
- 使用来自网格温度/速度探针的状态空间,将其输入到基于 MLP 的策略在 PPO RL 框架中。
- 将动作离散化为在10个分段上的分段常量温度轮廓,使用二进制水平并归一化以满足约束。
- 通过时间平均的 Nu 及其瞬时形式 Nu_inst 来评估性能。
实验结果
研究问题
- RQ1RL 基于控制在固定雷数下是否能比线性控制方法更有效地降低对流传热?
- RQ2在 RL 与线性控制下,临界雷数的可实现提升幅度是多少?
- RQ3控制延迟和可观测性如何影响在混沌区域中稳定或抑制 RB 的可行性?
- RQ4在 RL 控制下出现哪些流场结构导致热传输减少?
主要发现
- RL 控制将临界雷数从约 1e3(未受控)提高到约 1e4(线性)再到约 3e4(RL)。
- 对于 Ra > 3e4,RL 控制使时间平均 Nu 相对于未控制情形降低约 2.5,优于线性方法在 Ra < 1e6 时实现的约 1.5 的降低。
- RL 控制在 Ra 约为 3e4 时稳定导热态,在较高的 Ra 下实现稳定态或降低 Nu,而不是线性控制所典型的周期性流动。
- RL 诱导的流动结构类似于双环制,实质性地通过改变对流结构来降低热传输;这一效应可观测至 Ra 约 1e5,在 Ra 1e6 时虽减弱但仍存在。
- 训练时间随 Ra 而异,在 Ra ≲ 1e5 时在 V100 上不足一小时,在 Ra ≳ 1e6 时约为 150 小时,这是由于混沌性增加。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。