한양대학교 · 경제학
Jun Moon 교수의 연구실은 대규모 다중 에이전트 시스템의 최적화 및 강건성 제어를 중심으로, 리스크 민감도 기반의 스토크라스틱 게임 이론과 강화학습 기반 제어 기법을 융합한 연구를 수행합니다. 주요 연구 방향은 리스크-센티브 제어, 의미장(마이크로) 기반 게임 이론, 그리고 UAV 등에서의 경로 계획 및 충돌 회피 제어에 응용되는 딥 강화학습 알고리즘 개발입니다. 특히, 비선형 동적 시스템과 불확실성 환경에서의 안정적 제어 설계에 초점을 맞추고 있습니다.
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
This paper considers two classes of large population stochastic differential games connected to optimal and robust decentralized control of large-scale multiagent systems. The first problem (P1) is one where each agent minimizes an exponentiated cost function, capturing risk-sensitive behavior, whereas in the second problem (P2) each agent minimizes a worst-case risk-neutral cost function, where the “worst case” stems from the presence of an adversary entering each agent's dynamics characterized
In this paper, we propose a soft actor–critic (SAC) algorithm with hindsight experience replay (HER), called SACHER, which is a class of deep reinforcement learning (DRL) algorithm. SAC is an off-policy model-free DRL algorithm that outperforms earlier DRL algorithms in terms of exploration and robustness. However, in SAC, maximizing the entropy-augmented objective degrades the optimality of learning outcomes. We propose SACHER to improve the learning performance of SAC. We apply SACHER to the p
We consider two-player risk-sensitive zero-sum differential games (RSZSDGs). In our problem setup, both the drift term and the diffusion term in the controlled stochastic differential equation are dependent on the state and controls of both players, and the objective functional is of the risk-sensitive type. First, a stochastic maximum principle type necessary condition for an open-loop saddle point of the RSZSDG is established via nonlinear transformations of the adjoint processes of the equiva
In this note, we consider the linear-quadratic stochastic zero-sum differential game (LQ-SZSDG) for the Markov jump system (MJS) driven by Brownian motion. Unlike previous work considered in the literature, the diffusion term of the MJS is dependent on the state and the control of both players, and the cost parameters need not be definite matrices. We obtain a sufficient condition under which a feedback saddle point for the LQ-SZSDG exists. We show that the corresponding feedback saddle point is
In this article, we consider the linear-quadratic time-inconsistent mean-field type leader-follower Stackelberg differential game with an adapted open-loop information structure. The objective functionals of the leader and the follower include conditional expectations of state and control (mean field) variables, and the cost parameters could be general nonexponential discounting depending on the initial time. As stated in the existing literature, these two general settings of the objective funct
In this article, we consider the generalized risk-sensitive optimal control problem, where the objective functional is defined by the controlled backward stochastic differential equation (BSDE) with quadratic growth coefficient. We extend the earlier results of the risk-sensitive optimal control problem to the case of the objective functional given by the controlled BSDE. Note that the risk-neutral stochastic optimal control problem corresponds to the BSDE objective functional with linear growth
We consider robust stochastic large population games for coupled Markov jump linear systems (MJLSs). The N agents’ individual MJLSs are governed by different infinitesimal generators, and are affected not only by the control input but also by an individual disturbance (or adversarial) input. The mean field term, representing the average behaviour of N agents, is included in the individual worst-case cost function to capture coupling effects among agents. To circumvent the computational complexit
In this technical note, we consider linear exponential quadratic (LEQ) control for mean field stochastic differential equations (MFSDEs). The MFSDE includes the expectation value of state and control, and the objective functional is exponential of a quadratic functional in state, control, and their expectations. We obtain the explicit optimal solution as well as the optimal cost. The corresponding optimal solution is linear in state and its expectation, which is characterized by the Riccati diff
This paper considers linear-quadratic (LQ) stochastic leader-follower Stackelberg differential games for jump-diffusion systems with random coefficients. We first solve the LQ problem of the follower using the stochastic maximum principle and obtain the state-feedback representation of the open-loop optimal solution in terms of the integro-stochastic Riccati differential equation (ISRDE), where the state-feedback-type control is shown to be optimal via the completion of squares method. Next, we
We consider a class of stochastic differential games with the Stackelberg mode of play, with one leader and N uniform followers (where N is sufficiently large), where each player has its own local controlled dynamics and quadratic cost function, with the coupling between the players being through the cost functions. Particularly, the leader's cost function has as input the average value of the states of the followers, and each follower's cost function has a similar term in addition to being dire