[论文解读] Information Revelation and Alignment Faking in Stochastic Differential Games
本论文在部分信息条件下建立对称的两方线性二次随机微分博弈,提出对齐伪装控制以量化信息揭示,并分析可实施的基线、代理费舍信息以及伪装检测。给出半显式的基于 Riccati 的表征,并在模型错误指定下通过数值演示信息增益与可检测性的关系。
In competitive games with private objectives, actions can reveal information about hidden parameters. Quantifying such information revelation, however, is substantially more challenging, since it depends not only on the opponent's hidden parameter but also on the opponent's model of the game. We study this problem via a two-player linear-quadratic stochastic differential game under partial information, in which each player knows its own coupling parameter and models the opponent's hidden parameter through a prior. Starting from the full-information game, we characterize the Nash equilibrium by coupled Riccati equations. We then define baseline implementable controls by averaging the equilibrium under each player's prior. Building on this baseline, we formulate an alignment-faking control problem in which one player trades off fidelity to its implementable policy against information acquisition about the opponent's hidden parameter. The information incentive is constructed from a proxy Fisher information matrix based only on the player's available model. This leads to a tractable saddle-point formulation with semi-explicit control characterization through Riccati systems. Numerical illustrations show that alignment faking can substantially improve information gain over baseline play when the faker's model is accurate, but often at the cost of greater detectability. They also show that the proxy Fisher information can systematically differ from the true information under model misspecification.
研究动机与目标
- 量化对手隐藏参数在具有私有目标的两方随机微分博弈中的信息揭示。
- 表征通过对抗方对对手隐藏参数的先验平均得到的全信息纳什均衡的实现性基线控制。
- 引入对齐伪装控制问题,使信息增益与接近基线的可控性之间达到权衡。
- 开发基于代理费舍信息的目标并为 AF 控制构建可行的鞍点表述。
- 通过数值实验研究在模型错误指定下对齐伪装的检测及其影响。
提出的方法
- 在部分信息下建立对称的两方连续时间随机微分博弈,耦合的 Riccati 方程支配全信息纳什均衡。
- 将基线可实现控制定义为在每个玩家对对手隐藏参数的先验下,全信息均衡的期望。
- 为一方引入对齐伪装(AF)控制问题,在保真度与对对手隐藏参数信息获取之间进行权衡,利用一个代理费舍信息矩阵。
- 构造仅使用可获得量的代理 AF 目标,并形成一个可处理的极小极大鞍点问题。
- 通过 Riccati 系统得到半显式的控制表征,并结合基于 Riccati 的最小化和在辅助变量上的梯度步伐的迭代算法求解鞍点问题。
- 提供一个检测方案,即对手对残差与基线预测进行检验以检测 AF 行为。
- 在时间区间有界及确保 AF 动力学良定性的条件下,证明 Riccati 解的存在性与唯一性。

实验结果
研究问题
- RQ1Q1. 玩家关于对手隐藏参数的信息如何依赖于双方的先验 πA 和 πB?
- RQ2Q2. 对齐伪装控制是否在保持接近基线的同时增加对对手隐藏参数的信息增益?
- RQ3Q3. 当伪装者使用代理目标时,如何从观测轨迹中检测对齐伪装?
- RQ4Q4. 模型错误指定如何影响代理费舍信息及由此产生的 AF 策略?
主要发现
- 对齐伪装在伪装模型准确时能显著增加相对于基线的信息增益,但在实际执行中可能更易被检测到。
- 基线可实现控制通过在每方对对手隐藏参数的先验下对全信息纳什均衡进行平均得到。
- 构建了一个仅使用伪装者可获得量的代理费舍信息目标,使问题可解并通过半显式 Riccati 控制实现一个可处理的鞍点形式。
- 信息质量与 AF 的有效性主要取决于伪装者的模型,对手方的模型具有次要但显著的影响。
- 在模型错误指定下,代理费舍信息可能系统性偏离真实信息,影响 AF 策略及其感知效果。
![Figure 3: True asymptotic variance $[I(\gamma)^{-1}]_{m_{B},m_{B}}$ for $\mu_{A}\in\{1.0,1.25,1.5,1.75,2.0\}$ and $\mu_{B}\in\{1.0,1.25,1.5,1.75,2.0,2.25\}$ under both AF (solid) and no AF (dashed) gameplay. Parameters: $q^{AF}=5.0$ , $\lambda^{AF}=2.5$ , and $\rho_{A}=\rho_{B}=0.1$ .](https://ar5iv.labs.arxiv.org/html/2603.17197/assets/measure_info.png)
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。