[Paper Review] Information Revelation and Alignment Faking in Stochastic Differential Games
The paper develops a symmetric two-player linear-quadratic stochastic differential game under partial information, introduces alignment-faking controls to quantify information reveal, and analyzes implementable baselines, proxy Fisher information, and detection of faking. It provides semi-explicit Riccati-based characterizations and numerical demonstrations of information gain vs. detectability under model misspecification.
In competitive games with private objectives, actions can reveal information about hidden parameters. Quantifying such information revelation, however, is substantially more challenging, since it depends not only on the opponent's hidden parameter but also on the opponent's model of the game. We study this problem via a two-player linear-quadratic stochastic differential game under partial information, in which each player knows its own coupling parameter and models the opponent's hidden parameter through a prior. Starting from the full-information game, we characterize the Nash equilibrium by coupled Riccati equations. We then define baseline implementable controls by averaging the equilibrium under each player's prior. Building on this baseline, we formulate an alignment-faking control problem in which one player trades off fidelity to its implementable policy against information acquisition about the opponent's hidden parameter. The information incentive is constructed from a proxy Fisher information matrix based only on the player's available model. This leads to a tractable saddle-point formulation with semi-explicit control characterization through Riccati systems. Numerical illustrations show that alignment faking can substantially improve information gain over baseline play when the faker's model is accurate, but often at the cost of greater detectability. They also show that the proxy Fisher information can systematically differ from the true information under model misspecification.
Motivation & Objective
- Quantify information revelation about an opponent's hidden parameter in a two-player stochastic differential game with private objectives.
- Characterize implementable baseline controls obtained by averaging full-information Nash equilibria over priors.
- Introduce alignment-faking controls balancing information gain and proximity to baseline play.
- Develop a proxy Fisher information-based objective and a tractable saddle-point formulation for AF control.
- Investigate detection of alignment faking and effects of model misspecification through numerical experiments.
Proposed method
- Formulate a symmetric two-player continuous-time stochastic differential game under partial information with coupled Riccati equations governing the full-information Nash equilibrium.
- Define baseline implementable controls as expectations of the full-information equilibrium under each player's prior on the opponent's hidden parameter.
- Introduce an alignment-faking (AF) control problem for one player to trade fidelity to the baseline against information acquisition about the opponent's hidden parameter, using a proxy Fisher information matrix.
- Construct a proxy AF objective that uses only available quantities and formulate a max-min saddle-point problem.
- Obtain semi-explicit control characterizations via Riccati systems and solve the saddle-point problem with an iterative algorithm combining Riccati-based minimization and gradient steps in auxiliary variables.
- Provide a detection scheme where the opponent tests residuals against baseline predictions to detect AF behavior.
- Prove existence and uniqueness of the Riccati solutions under time horizon bounds and under conditions ensuring well-posedness of AF dynamics.

Experimental results
Research questions
- RQ1Q1. How does a player's information about the opponent's hidden parameter depend on the players' priors πA and πB?
- RQ2Q2. Can an alignment-faking control increase information gain about the opponent's hidden parameter while remaining close to baseline play?
- RQ3Q3. How can one detect alignment fakings from observed trajectories when the faker uses a proxy objective?
- RQ4Q4. How does model misspecification affect proxy Fisher information and the resulting AF strategy?
Key findings
- Alignment faking can substantially increase information gain over baseline play when the faker's model is accurate, but may become more detectable in practice.
- Baseline implementable controls are obtained by averaging the full-information Nash equilibrium under each player's prior on the opponent's hidden parameter.
- A proxy Fisher information-based objective is constructed using quantities available to the faker, enabling a tractable saddle-point formulation with semi-explicit Riccati-based controls.
- The information quality and the effectiveness of AF depend mainly on the faker's model, with the opponent's model having a secondary but noticeable effect.
- Proxy Fisher information can systematically diverge from true information under model misspecification, affecting the AF strategy and its perceived effectiveness.
![Figure 3: True asymptotic variance $[I(\gamma)^{-1}]_{m_{B},m_{B}}$ for $\mu_{A}\in\{1.0,1.25,1.5,1.75,2.0\}$ and $\mu_{B}\in\{1.0,1.25,1.5,1.75,2.0,2.25\}$ under both AF (solid) and no AF (dashed) gameplay. Parameters: $q^{AF}=5.0$ , $\lambda^{AF}=2.5$ , and $\rho_{A}=\rho_{B}=0.1$ .](https://ar5iv.labs.arxiv.org/html/2603.17197/assets/measure_info.png)
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.