[論文レビュー] Data-based approximate policy iteration for nonlinear continuous-time optimal control design
本稿では、システムモデルを必要とせずに非線形連続時刻最適制御を実現するデータベース型の近似政策反復(API)手法を提案する。アクター・クリティック型ニューラルネットワークと最小二乗加重残差法を用い、実システムのデータから直接ハミルトニアン=ジャコビ=ベルヌーイ方程式の解および最適制御方策を学習し、制約なしおよび入力制約付きの両状況で収束性と有効な制御を達成した。非線形ベンチマーク系(入力制限が厳しい)でも検証された。
This paper addresses the model-free nonlinear optimal problem with generalized cost functional, and a data-based reinforcement learning technique is developed. It is known that the nonlinear optimal control problem relies on the solution of the Hamilton-Jacobi-Bellman (HJB) equation, which is a nonlinear partial differential equation that is generally impossible to be solved analytically. Even worse, most of practical systems are too complicated to establish their accurate mathematical model. To overcome these difficulties, we propose a data-based approximate policy iteration (API) method by using real system data rather than system model. Firstly, a model-free policy iteration algorithm is derived for constrained optimal control problem and its convergence is proved, which can learn the solution of HJB equation and optimal control policy without requiring any knowledge of system mathematical model. The implementation of the algorithm is based on the thought of actor-critic structure, where actor and critic neural networks (NNs) are employed to approximate the control policy and cost function, respectively. To update the weights of actor and critic NNs, a least-square approach is developed based on the method of weighted residuals. The whole data-based API method includes two parts, where the first part is implemented online to collect real system information, and the second part is conducting offline policy iteration to learn the solution of HJB equation and the control policy. Then, the data-based API algorithm is simplified for solving unconstrained optimal control problem of nonlinear and linear systems. Finally, we test the efficiency of the data-based API control design method on a simple nonlinear system, and further apply it to a rotational/translational actuator system. The simulation results demonstrate the effectiveness of the proposed method.
研究の動機と目的
- 正確なシステムモデルが入手できない状況における非線形連続時刻最適制御の課題に対処すること。
- システムのデータから直接最適制御方策を学習するモデルフリーな強化学習手法を開発すること。
- ニューラルネットワーク近似を用いて入力制約下でも政策反復アルゴリズムの収束を保証すること。
- 入力飽和を伴う複雑な回転・直線運動アクチュエータ(RTAC)ベンチマークのような複雑な非線形系に対しても有効性を示すこと。
- 従来のモデル依存型最適制御設計に対する実用的でデータベース型の代替手法を提供すること。
提案手法
- オンラインデータ収集とオフライン政策反復の2段階で動作するデータベース型の近似政策反復(API)アルゴリズムを開発した。
- アクター・クリティック型ニューラルネットワーク構造を採用し、クリティックがコスト・トゥ・ゴール関数を近似し、アクターが制御方策を近似する。
- システムダイナミクスを必要とせず、重み付き残差法に基づく最小二乗法を用いてニューラルネットワークの重みを更新する。
- 非線形PDEの解析的解法を回避するため、実システムの軌道を用いて一般化ハミルトニアン=ジャコビ=ベルヌーイ(GHJB)方程式を繰り返し解く。
- 入力制約は、双曲正接関数を用いた制御入力の非線形変換により処理する。
- 提案されたデータ駆動型フレームワーク下で、アルゴリズムの収束性を理論的に証明した。
実験結果
リサーチクエスチョン
- RQ1モデルフリーな政策反復手法が、システムモデルを必要とせずに非線形連続時刻最適制御問題を効果的に解けるか?
- RQ2ニューラルネットワークを用いて、データベース型最適制御設計にどのようにして入力制約を統合できるか?
- RQ3提案されたデータベース型API手法が、実システムのデータのみを用いて最適解に収束するか?
- RQ4入力飽和を伴う複雑な非線形系(例:RTACベンチマーク)の制御において、この手法はどの程度有効か?
- RQ5システムダイナミクスを必要としない最小二乗加重残差法が、正確にニューラルネットワークの重みを更新できるか?
主な発見
- データベース型APIアルゴリズムは、制約なしおよび制約ありの非線形系の両方において、実システムのデータのみを用いて最適解に収束した。
- 入力制約 |u| ≤ 0.2 を有するRTAC系では、最終的な制御方策が制約を満たし、シミュレーション全体を通して制御入力が境界内に保たれた。
- 時間の経過に伴い、コスト関数 J(t) は 0.6781 に収束し、学習済み方策による有効な最適性能を示した。
- クリティックおよびアクターのニューラルネットワーク重みベクトルのノルムは、それぞれ 8.1462 および 31.7143 に収束し、学習の安定性を示した。
- 図15~18に示すように、代表的な最初の6つのクリティックおよびアクターNN重みが反復回数とともに収束し、学習の進行を確認した。
- 本手法は、単純な非線形系および複雑なRTACベンチマークの両方で優れた性能を示し、実用的応用性を実証した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。