Skip to main content
QUICK REVIEW

[논문 리뷰] Deep Learning for Mean Field Optimal Transport

Sebastian Baudelet, Brieuc Frénais|arXiv (Cornell University)|2023. 02. 28.
Climate Change Policy and EconomicsEconomics, Econometrics and Finance인용 수 3
한 줄 요약

이 논문은 에이전트가 협력적으로 사회적 비용을 최소화하면서 고정된 종단 분포에 도달하는 평균장 최적운송(MFOT) 문제를 해결하기 위한 세 가지 딥러닝 기반 수치적 방법을 제안한다. 이 방법들은 신경망을 활용해 최적의 제어를 학습하거나, 전진-후향 PDE 시스템을 해결하거나, 확장된 라그랑주 원-이중 프레임워크를 사용하며, 선형-제곱형 및 혼잡도 모델링 테스트 케이스에서 정확한 해를 보여준다.

ABSTRACT

Mean field control (MFC) problems have been introduced to study social optima in very large populations of strategic agents. The main idea is to consider an infinite population and to simplify the analysis by using a mean field approximation. These problems can also be viewed as optimal control problems for McKean-Vlasov dynamics. They have found applications in a wide range of fields, from economics and finance to social sciences and engineering. Usually, the goal for the agents is to minimize a total cost which consists in the integral of a running cost plus a terminal cost. In this work, we consider MFC problems in which there is no terminal cost but, instead, the terminal distribution is prescribed. We call such problems mean field optimal transport problems since they can be viewed as a generalization of classical optimal transport problems when mean field interactions occur in the dynamics or the running cost function. We propose three numerical methods based on neural networks. The first one is based on directly learning an optimal control. The second one amounts to solve a forward-backward PDE system characterizing the solution. The third one relies on a primal-dual approach. We illustrate these methods with numerical experiments conducted on two families of examples.

연구 동기 및 목표

  • 종단 분포가 고정된 평균장 최적운송 문제를 다루되, 종단 비용을 최소화하는 것이 아니라.
  • 스콜레링 브리지와 표준 MFG/MFC 설정을 초월하는 일반화된 MFOT 문제를 해결하기 위한 딥러닝 기반 수치적 방법을 개발한다.
  • 복잡한 평균장 상호작용과 비트리비얼한 동역학을 가진 문제를 해결할 수 있도록 기존 방법의 한계를 극복한다.
  • 고차원 환경에서 최적의 제어와 밀도를 근사하기 위해 신경망을 사용한 확장 가능하고 미분 가능한 접근법을 제공한다.
  • 벌점 항, PDE 제약 조건, 확장된 라그랑주 공식화를 통한 손실 기반 훈련을 통해 수렴성과 정확성을 확보한다.

제안 방법

  • Method 1은 딥 강화학습 유사 접근법을 사용한다: 신경망이 매킨-브라운 확률미분방정식의 몬테카를로 롤아웃을 통해 최적의 제어 정책을 학습하며, 종단 분포의 편차에 대한 벌점이 부여된다.
  • Method 2는 MFOT 문제의 최적성 조건에서 유도된 전진-후향 PDE 시스템의 해를 신경망을 통해 직접 근사한다.
  • Method 3은 종단 분포 제약 조건을 원-이중 최적화 기반으로 강제하기 위해 확장된 라그랑주 공식을 사용하며, 가치 함수, 밀도, 라그랑주 승수에 대해 별도의 네트워크를 사용한다.
  • 모든 방법은 완전히 연결된 피드포워드 신경망을 사용하며, 잔차 연결, ReLU 또는 시그모이드 활성화 함수를 적용하고, 미니배치 샘플링을 통한 확률적 경사하강법으로 훈련된다.
  • 확장된 라그랑주에서의 펜alty 가중치 $ C_W $, $ C_0^{(KFP)} $, $ C_T^{(KFP)} $, $ C^{(KFP)} $, $ C^{(HJB)} $ 및 $ r $ 등의 하이퍼파rameter는 안정성과 수렴성을 확보하기 위해 경험적으로 조정된다.
  • LQ 문제의 경우, 정확도 향상과 분석적 해와의 일치를 위해 제어 네트워크 출력에 추가적인 제곱형 보정 항이 삽입된다.
Figure 1 : Evolution of the density in the LQ Test case 1. Each plot corresponds to one time step and displays the densities as functions of the space variable, $\textstyle x$ . The densities are: The density obtained by applying the control learnt by each of the three deep learning methods as well
Figure 1 : Evolution of the density in the LQ Test case 1. Each plot corresponds to one time step and displays the densities as functions of the space variable, $\textstyle x$ . The densities are: The density obtained by applying the control learnt by each of the three deep learning methods as well

실험 결과

연구 질문

  • RQ1딥러닝은 종단 비용이 없고 종단 분포가 고정된 평균장 최적운송 문제를 효과적으로 해결할 수 있는가?
  • RQ2MFOT 문제의 수렴성과 정확성 측면에서 다양한 신경망 아키텍처와 손실 공식화 간의 성능 비교는 어떻게 되는가?
  • RQ3제안된 방법들은 비선형 동역학과 혼잡도 효과와 같은 복잡한 평균장 상호작용을 처리할 수 있는가?
  • RQ4벌점 가중치와 확장된 라그랑주 매개변수 $ r $ 와 같은 하이퍼파rameter는 훈련 안정성과 성능에 어떤 영향을 미치는가?
  • RQ5딥러닝 기반 방법은 스콜레링 브리지와 선형-제곱형 케이스를 초월해 얼마나 일반화될 수 있는가?

주요 결과

  • 제안된 세 가지 방법은 선형-제곱형(LQ) 테스트 케이스에서 분석적 해와 높은 정확도로 일치시키며, 수렴성과 강건성을 입증한다.
  • Method 1은 벌점 항 $ C_W $ 를 상태 차원과 계산 비용에 따라 조정함으로써 최적의 제어 정책을 성공적으로 학습한다.
  • Method 2는 손실 항을 통해 초기 조건, 종단 조건, PDE 제약 조건을 강제함으로써 전진-후향 PDE 시스템을 신경망을 통해 효과적으로 해결한다.
  • 확장된 라그랑주 접근법을 사용한 Method 3는 $ r = 0.1 $ 일 때 안정적인 수렴을 보이며, 라그랑주 승수 네트워크에 시그모이드 활성화 함수를 사용함으로써 유한한 밀도 추정치를 확보한다.
  • 모든 방법은 혼잡도 효과를 포함한 비트리비얼한 평균장 상호작용을 성공적으로 처리한다. 이는 혼잡도 테스트 케이스에서 입증되었다.
  • 수치 실험을 통해 방법들이 고차원 환경에서도 확장 가능하고 효과적임을 확인하였으며, 하이퍼파rameter의 세심한 조정과 미니배치 샘플링을 통해 훈련이 안정화됨을 입증하였다.
Figure 2 : Evolution of the control in the LQ Test case 1. Each plot corresponds to one time step and displays the controls as functions of the space variable, $\textstyle x$ . The controls are: The control learnt by each of the three deep learning methods as well as the ground-truth control given b
Figure 2 : Evolution of the control in the LQ Test case 1. Each plot corresponds to one time step and displays the controls as functions of the space variable, $\textstyle x$ . The controls are: The control learnt by each of the three deep learning methods as well as the ground-truth control given b

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.