[论文解读] Inverse Active Sensing: Modeling and Understanding Timely Decision-Making
本文提出了一种统一的反向主动感知框架,用于建模在内生性、情境依赖的时间压力下及时决策的问题,平衡准确性、成本与速度。该框架通过贝叶斯反向推理方法,从观测行为中推断潜在偏好与策略,展示了如何量化医疗诊断和招聘等现实决策过程中的权衡。
Evidence-based decision-making entails collecting (costly) observations about an underlying phenomenon of interest, and subsequently committing to an (informed) decision on the basis of accumulated evidence. In this setting, active sensing is the goal-oriented problem of efficiently selecting which acquisitions to make, and when and what decision to settle on. As its complement, inverse active sensing seeks to uncover an agent's preferences and strategy given their observable decision-making behavior. In this paper, we develop an expressive, unified framework for the general setting of evidence-based decision-making under endogenous, context-dependent time pressure---which requires negotiating (subjective) tradeoffs between accuracy, speediness, and cost of information. Using this language, we demonstrate how it enables modeling intuitive notions of surprise, suspense, and optimality in decision strategies (the forward problem). Finally, we illustrate how this formulation enables understanding decision-making behavior by quantifying preferences implicit in observed decision strategies (the inverse problem).
研究动机与目标
- 为解决现有主动感知模型表达能力不足且需要完全指定主观偏好的局限性。
- 建模在内生性与情境依赖时间压力下的决策过程,其中截止时间与信息成本取决于代理的选择与情境。
- 开发一种反向主动感知方法——从观测到的决策行为中推断代理的隐性偏好与策略。
- 在单一概率框架内统一正向问题(建模最优策略)与反向问题(重构偏好)。
- 通过量化实时决策系统中准确性、成本与速度的权衡,使该框架在医学、招聘与认知科学等领域的实际应用成为可能。
提出的方法
- 提出一个统一的概率框架,将决策过程建模为状态依赖的时间压力序列过程,整合成本、截止时间与准确性惩罚。
- 使用贝叶斯识别模型表示信念更新与策略选择,将时间压力建模为情境依赖的内生过程。
- 引入对参数(η, ρ, κ)的层次先验,以建模偏好、成本与策略动态中的不确定性。
- 使用贝叶斯规则在策略上定义后验分布,其似然由给定策略下观测到的决策路径的概率定义。
- 通过边缘化潜变量推导后验分布,动态项相互抵消,从而获得可处理的推理过程。
- 在温和正则性条件下,证明后验分布对参数η与ρ具有可微性,从而支持基于梯度的推理。
实验结果
研究问题
- RQ1我们如何建模能够考虑决策过程中内生性与情境依赖时间压力的主动感知策略?
- RQ2主观上对准确性、成本与速度的权衡在何种程度上塑造了最优决策策略?
- RQ3我们如何从观测行为中推断代理的潜在偏好与决策策略?
- RQ4在所提出的框架中,惊讶与悬念以何种方式自然涌现?
- RQ5该框架能否用于在现实应用中反向工程决策行为,如医疗诊断或招聘?
主要发现
- 该框架成功将惊讶与悬念等直观概念建模为在时间压力下信念动态的自然结果。
- 策略的后验分布几乎处处对偏好参数η与ρ可微,从而支持基于梯度的推理。
- 该方法可在无需显式指定惩罚或成本的情况下实现偏好的反向推理,克服了先前模型的关键局限。
- 该框架通过同时整合内生性时间压力与情境依赖动态,推广了现有模型,从而能够更丰富地建模现实世界的决策系统。
- 实证验证表明,该模型能够从序列决策任务中的观测行为中准确重构决策策略与潜在偏好。
- 该方法能够量化决策准确性、信息获取成本与截止时间违规之间的权衡,为系统设计提供可操作的洞见。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。