[论文解读] Counterfactual Explanations in Sequential Decision Making Under Uncertainty
本文提出了一种新颖的框架,用于在不确定性下的序列决策过程中生成反事实解释,采用有限时域马尔可夫决策过程和Gumbel-Max结构因果模型。该方法提出了一种多项式时间动态规划算法,可计算出最优反事实策略,通过最多改变$k$个动作来保证结果改善,已在合成数据和真实认知行为疗法数据上验证,提供了具有临床意义的洞察。
Methods to find counterfactual explanations have predominantly focused on one step decision making processes. In this work, we initiate the development of methods to find counterfactual explanations for decision making processes in which multiple, dependent actions are taken sequentially over time. We start by formally characterizing a sequence of actions and states using finite horizon Markov decision processes and the Gumbel-Max structural causal model. Building upon this characterization, we formally state the problem of finding counterfactual explanations for sequential decision making processes. In our problem formulation, the counterfactual explanation specifies an alternative sequence of actions differing in at most k actions from the observed sequence that could have led the observed process realization to a better outcome. Then, we introduce a polynomial time algorithm based on dynamic programming to build a counterfactual policy that is guaranteed to always provide the optimal counterfactual explanation on every possible realization of the counterfactual environment dynamics. We validate our algorithm using both synthetic and real data from cognitive behavioral therapy and show that the counterfactual explanations our algorithm finds can provide valuable insights to enhance sequential decision making under uncertainty.
研究动机与目标
- 填补在具有依赖动作的多步序列决策过程中反事实解释的空白。
- 将反事实解释形式化为与观察序列最多相差$k$个动作的行动序列上的约束搜索。
- 开发一种可扩展的、多项式时间的算法,确保在所有反事实动态实现中获得最优的反事实结果。
- 在真实世界临床数据上验证该方法,展示可操作且具有临床可解释性的洞察。
- 为回顾性分析序列决策结果提供一种基于干预的因果解释框架。
提出的方法
- 使用具有离散、低维状态和动作空间的有限时域马尔可夫决策过程(MDPs)来形式化序列决策过程。
- 使用Gumbel-Max结构因果模型来建模转移概率,以确保反事实稳定性并实现可靠的后果估计。
- 将反事实解释问题定义为与观察序列最多相差$k$个动作的替代行动序列上的约束搜索。
- 开发一种基于动态规划的算法,可计算出反事实策略,保证在所有可能的反事实转移动态实现中产生最优结果。
- 采用策略评估机制,通过采样多个反事实实现来评估结果改善情况,并识别关键干预时间点。
- 使用算法采样方法识别在多个反事实实现中一致的动作变更(例如,用行为激活替代认知重构)。
实验结果
研究问题
- RQ1反事实解释能否有意义地扩展到具有多个依赖动作的序列决策过程?
- RQ2如何正式定义并计算反事实解释,以确保在不确定性下获得最优结果?
- RQ3寻找与观察路径最多相差$k$个动作的最优反事实行动序列的计算复杂度是多少?
- RQ4从所提方法得出的反事实解释与真实世界序列决策过程中的观察结果相比如何?
- RQ5在不同实现中,最优反事实策略最常推荐哪些时间点和动作变更?
主要发现
- 对于抑郁症状恶化的患者,当$k=3$时,最优反事实策略使平均结果相比观察结果改善了9.5%。
- 在85%的采样反事实实现中,结果优于观察结果,表明在建议的动作变更下,改善的可能性很高。
- 时间点$t=10$、$t=13$和$t=16$在反事实解释中被显著突出,其中$t=10$标志着临床恶化的开始。
- 最优策略始终一致地建议在恶化阶段开始时,将认知重构技术(CRT)替换为行为激活(BHA),这一建议得到了心理健康专家的认可,具有临床合理性。
- 在表现最佳的反事实实现中,恶化轨迹被很大程度上避免,表明该方法具有预防负面结果的潜力。
- 该方法识别出与临床最佳实践一致的可操作、可解释的治疗序列变更,增强了在现实应用中的信任度和实用性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。