Skip to main content
QUICK REVIEW

[論文レビュー] Deep reinforcement learning for the control of conjugate heat transfer with application to workpiece cooling

Elie Hachem, Hassan Ghraieb|arXiv (Cornell University)|Nov 30, 2020
Heat Transfer Mechanisms参考文献 65被引用数 4
ひとこと要約

本稿では、流体・構造系における共役熱伝達を最適化する目的で、退化した近接方策最適化(PPO)アルゴリズムを用いた新規な深層強化学習(DRL)フレームワークを提案する。2次元および3次元の設定において、自然対流および強制対流の制御が有効に行われ、温度均一性が向上し、対称的駆動下での非対称なワークピece配置といった直感的でない最適な配置が同定された。

ABSTRACT

This research gauges the ability of deep reinforcement learning (DRL) techniques to assist the control of conjugate heat transfer systems governed by the coupled Navier--Stokes and heat equations. It uses a novel, "degenerate" version of the proximal policy optimization (PPO) algorithm, intended for situations where the optimal policy to be learnt by a neural network does not depend on state, as is notably the case in optimization and open-loop control problems. The numerical reward fed to the neural network is computed with an in-house stabilized finite elements environment combining variational multi-scale (VMS) modeling of the governing equations, immerse volume method, and multi-component anisotropic mesh adaptation. Several test cases of natural and forced convection in two and three dimensions are used as testbed for developing the methodology. The approach successfully alleviates the natural convection induced enhancement of heat transfer in a two-dimensional, differentially heated square cavity controlled by piece-wise constant fluctuations of the sidewall temperature. It also proves capable of improving the homogeneity of temperature across the surface of two and three-dimensional hot workpieces under impingement cooling. Various cases are tackled, in which the position of multiple cold air injectors is optimized relative to a fixed workpiece position. The flexibility of the numerical framework makes it tractable to solve also the inverse problem, i.e., to optimize the workpiece position relative to a fixed injector distribution. The obtained results showcase the potential of the method for black-box optimization of practically meaningful computational fluid dynamics (CFD) conjugate heat transfer systems.

研究の動機と目的

  • 連成ナビエ–ストークス方程式および熱伝導方程式に従う共役熱伝達の制御戦略をDRLベースで開発すること。
  • 事前の知識が限られる状況下で、大規模かつ高次元のパrameter空間を最適化する課題に対処すること。
  • DRLが、従来の設計直感を超える予期せぬ高パフォーマンス制御配置を発見する可能性を検討すること。
  • 熱管理分野における前向き制御問題と逆問題の両方を解くことの柔軟性を示すこと。
  • 実際の2次元および3次元の共役熱伝達シナリオ、特に衝突冷却および温度勾配のあるキャビティを含む、妥当な設定で手法を検証すること。

提案手法

  • 状態依存のない行動選択が可能なため、オープンループおよび最適化問題に適した、ポリシー・ネットワークに状態依存性を持たない「退化した」PPOの変種を採用する。
  • 変分マルチスケール(VMS)モデル、埋め込まれた境界法、および異方的メッシュ適応を組み合わせたカスタム安定化有限要素ソルバを採用し、高精度なCFDシミュレーションを実現する。
  • 温度均一性および熱勾配に基づいて数値的に計算された報酬信号を用い、ポリシー学習を誘導する。
  • 複雑な幾何形状における解の精度向上と計算コスト低減を図るため、複数成分の異方的メッシュ適応戦略を統合する。
  • 同じフレームワーク内で、インジェクタ位置の最適化(前向き制御)とワークピース位置の最適化(逆問題制御)の両方を可能にする。
  • 深層ニューラルネットワークを用いて、シミュレーションデータから直接制御ポリシーを学習し、システムをブラックボックス最適化問題として扱う。

実験結果

リサーチクエスチョン

  • RQ1状態依存の行動選択が行えない状況下でも、退化したPPOアルゴリズムは共役熱伝達系において最適な制御ポリシーを効果的に学習できるか?
  • RQ2衝突冷却下での2次元および3次元の高温ワークピースにおいて、DRLは温度均一性をどの程度向上できるか?
  • RQ3DRLは、対称的駆動下での非対称なワークピース配置といった、直感的でないあるいは反直感的な最適配置を明らかにできるか?
  • RQ4工業的熱制御問題に一般的に見られる高次元パrameter空間を処理するにあたり、DRLフレームワークのスケーラビリティおよびロバスト性はどの程度か?
  • RQ5同じフレームワークで、共役熱伝達における前向き問題と逆問題の両方を効率的に解くことができるか?

主な発見

  • 退化したPPOアルゴリズムは、2次元の温度勾配を持つキャビティにおいて、自然対流に起因する熱伝達強化を、側壁温度のフラクチュエーション最適化によって効果的に抑制した。
  • 衝突冷却下の2次元および3次元の高温ワークピースにおいて、冷気インジェクタの空間的配置最適化によって、表面における温度均一性が顕著に向上した。
  • 逆問題設定において、DRLフレームワークは、対称的駆動下でも最適なワークピース位置が中心からずれていることを同定した。これは、対称性に基づく設計では予想できなかった直感に反する結果であった。
  • 本手法は、収束性および解の品質の面で、単純なパラメトリックスイープを上回る、高次元制御空間におけるロバスト性と効率性を示した。
  • フレームワークは、前向き制御問題と逆問題の両方を同じ数値環境で処理できることを示し、より広範な設計探索が可能になった。
  • 結果から、DRLは、従来の最適化手法やヒューリスティック設計手法では到達できない、新規で高パフォーマンスな制御戦略を発見できる可能性があることが示唆された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。