Skip to main content
QUICK REVIEW

[논문 리뷰] Information Revelation and Alignment Faking in Stochastic Differential Games

Daniel Ralston, Xu Yang|arXiv (Cornell University)|2026. 03. 17.
Game Theory and Applications인용 수 0
한 줄 요약

본 논문은 부분 정보 하에서 대칭적인 두 선수의 선형-이차 확률적 미분 게임을 개발하고, 정보 공개를 정량하기 위한 alignment-faking 컨트롤을 도입하며, 구현 가능한 baselines, 프록시 Fisher 정보, 그리고 위조 탐지를 분석합니다. 모형 오해추정 하에서 정보 이득 대 탐지 가능성의 반정식 Riccati 기반 특성과 수치적 시연을 제공합니다.

ABSTRACT

In competitive games with private objectives, actions can reveal information about hidden parameters. Quantifying such information revelation, however, is substantially more challenging, since it depends not only on the opponent's hidden parameter but also on the opponent's model of the game. We study this problem via a two-player linear-quadratic stochastic differential game under partial information, in which each player knows its own coupling parameter and models the opponent's hidden parameter through a prior. Starting from the full-information game, we characterize the Nash equilibrium by coupled Riccati equations. We then define baseline implementable controls by averaging the equilibrium under each player's prior. Building on this baseline, we formulate an alignment-faking control problem in which one player trades off fidelity to its implementable policy against information acquisition about the opponent's hidden parameter. The information incentive is constructed from a proxy Fisher information matrix based only on the player's available model. This leads to a tractable saddle-point formulation with semi-explicit control characterization through Riccati systems. Numerical illustrations show that alignment faking can substantially improve information gain over baseline play when the faker's model is accurate, but often at the cost of greater detectability. They also show that the proxy Fisher information can systematically differ from the true information under model misspecification.

연구 동기 및 목표

  • 상대의 숨겨진 매개변수에 대한 정보를 두 선수 간의 부분 정보 확률적 미분 게임에서 공개 정도를 정량화한다.
  • 각 플레이어의 상대의 숨겨진 매개변수에 대한 사전분포에 따라 전체 정보가 완전정보 Nash 균형으로부터 기대값으로 도출되는 구현 가능한 baseline 제어를 특징짓는다.
  • alignment-faking( AF) 제어 문제를 도입하여 baseline에 대한 충실도와 상대의 숨겨진 매개변수에 대한 정보 획득 사이의 균형을 맞춘다.
  • 프록시 Fisher 정보 기반 목표를 구성하고 AF 제어를 위한 계산 가능한 saddle-point 형식을 도입한다.
  • Riccati 시스템을 이용한 반정식적 제어 특성을 얻고, Riccati 기반의 최소화와 보조 변수의 기울기 단계가 결합된 반복 알고리즘으로 saddle-point 문제를 풀 수 있다.
  • 상대가 baseline 예측에 대해 잔차를 테스트하는 탐지 스킴을 제공한다.
  • 시간 구간의 경계 및 AF 다이나믹스의 잘 정의성을 보장하는 조건하에서 Riccati 해의 존재성 및 고유성를 증명한다.

제안 방법

  • 부분 정보 하에서 결합된 Riccati 방정식이 지배하는 완전 정보 Nash 균형을 평가하는 대칭 두 선수 연속시간 확률적 미분 게임을 형식화한다.
  • baseline 구현 가능한 제어는 상대의 숨겨진 매개변수에 대한 각 플레이어의 사전 하에 완전 정보 균형의 기댓값으로 정의한다.
  • 한 선수의 alignment-faking(AF) 제어 문제를 도입하여 baseline에 대한 충실도와 상대의 숨겨진 매개변수에 대한 정보 획득 사이를 거래하도록 한다(프록시 Fisher 정보 행렬을 사용).
  • 사용 가능한 양만을 이용하는 프록시 AF 목표를 구성하고 이를 이용해 최대-최소의 saddle-point 문제를 정식화한다.
  • Riccati 시스템을 통한 반정식적 제어 특성화를 얻고, Riccati 기반 최소화와 보조 변수의 기울기 단계가 결합된 반복 알고리즘으로 saddle-point 문제를 풀이한다.
  • 상대가 AF 행동을 탐지하기 위해 baseline 예측에 대한 잔차를 테스트하는 탐지 체계를 제공한다.
  • AF 다이나믹스의 잘 정의성을 보장하는 조건 하에서 시간 구간 경계 하 존재성 및 고유성을 증명한다.
Figure 2: Combined view of AF behavior with fixed $\mu_{A}=m_{A}=1.0$ and varying $\mu_{B}\in\{1.0,1.2,1.5,1.8,2.0,2.2\}$ . Top panels show state (left) and control (right) trajectories in the case that $\mu_{B}=m_{B}=1.0$ . Bottom panels show asymptotic variance (left) and regression-based detectab
Figure 2: Combined view of AF behavior with fixed $\mu_{A}=m_{A}=1.0$ and varying $\mu_{B}\in\{1.0,1.2,1.5,1.8,2.0,2.2\}$ . Top panels show state (left) and control (right) trajectories in the case that $\mu_{B}=m_{B}=1.0$ . Bottom panels show asymptotic variance (left) and regression-based detectab

실험 결과

연구 질문

  • RQ1Q1. 상대의 숨겨진 매개변수에 대한 플레이어의 정보가 두 플레이어의 사전 πA 및 πB에 어떻게 의존하는가?
  • RQ2Q2. alignment-faking 제어가 baseline 플레이에 근접한 상태를 유지하면서 상대의 숨겨진 매개변수에 대한 정보 이득을 증가시킬 수 있는가?
  • RQ3Q3. faker가 프록시 목표를 사용할 때 관찰된 궤적에서 alignment fakings를 어떻게 탐지할 수 있는가?
  • RQ4Q4. 모델 오오적성이 프록시 Fisher 정보와 그에 따른 AF 전략에 어떤 영향을 주는가?

주요 결과

  • Alignment faking은 faker의 모델이 정확할 때 baseline 대비 정보 이득을 크게 증가시킬 수 있지만, 실제로는 더 탐지하기 쉬워질 수 있다.
  • baseline 구현 가능한 제어는 상대의 숨겨진 매개변수에 대한 각 플레이어의 사전 하에 완전 정보 Nash 균형을 평균화하여 얻는다.
  • 프록시 Fisher 정보 기반 목표가 faker가 이용할 수 있는 양만으로 구성되어 계산 가능한 saddle-point 형식을 가능하게 하며 반정식 Riccati 기반 제어를 제공한다.
  • 정보의 품질과 AF의 효과는 주로 faker의 모델에 크게 의존하며 상대의 모델은 보조적이지만 뚜렷한 영향을 준다.
  • 모형 오오적성하에서 프록시 Fisher 정보가 실제 정보와 체계적으로 다르게 발산할 수 있어 AF 전략과 그 인식된 효과에 영향을 준다.
Figure 3: True asymptotic variance $[I(\gamma)^{-1}]_{m_{B},m_{B}}$ for $\mu_{A}\in\{1.0,1.25,1.5,1.75,2.0\}$ and $\mu_{B}\in\{1.0,1.25,1.5,1.75,2.0,2.25\}$ under both AF (solid) and no AF (dashed) gameplay. Parameters: $q^{AF}=5.0$ , $\lambda^{AF}=2.5$ , and $\rho_{A}=\rho_{B}=0.1$ .
Figure 3: True asymptotic variance $[I(\gamma)^{-1}]_{m_{B},m_{B}}$ for $\mu_{A}\in\{1.0,1.25,1.5,1.75,2.0\}$ and $\mu_{B}\in\{1.0,1.25,1.5,1.75,2.0,2.25\}$ under both AF (solid) and no AF (dashed) gameplay. Parameters: $q^{AF}=5.0$ , $\lambda^{AF}=2.5$ , and $\rho_{A}=\rho_{B}=0.1$ .

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.