Skip to main content
QUICK REVIEW

[论文解读] Balancing Suspense and Surprise: Timely Decision Making with Endogenous Information Acquisition

Ahmed M. Alaa, Mihaela van der Schaar|arXiv (Cornell University)|Oct 24, 2016
Auction Theory and Applications被引用 6
一句话总结

本文提出一个贝叶斯模型,用于在时间压力下进行及时决策,其中决策者内生地获取信息,通过从随机时间序列中采样来预测负面事件。最优策略表现出一种‘会合’结构,平衡了意外(信息增益)与悬念(生存风险),停止与继续区域取决于信念和当前情境(时间序列的当前实现)。

ABSTRACT

We develop a Bayesian model for decision-making under time pressure with endogenous information acquisition. In our model, the decision maker decides when to observe (costly) information by sampling an underlying continuous-time stochastic process (time series) that conveys information about the potential occurrence or non-occurrence of an adverse event which will terminate the decision-making process. In her attempt to predict the occurrence of the adverse event, the decision-maker follows a policy that determines when to acquire information from the time series (continuation), and when to stop acquiring information and make a final prediction (stopping). We show that the optimal policy has a rendezvous structure, i.e. a structure in which whenever a new information sample is gathered from the time series, the optimal "date" for acquiring the next sample becomes computable. The optimal interval between two information samples balances a trade-off between the decision maker's surprise, i.e. the drift in her posterior belief after observing new information, and suspense, i.e. the probability that the adverse event occurs in the time interval between two information samples. Moreover, we characterize the continuation and stopping regions in the decision-maker's state-space, and show that they depend not only on the decision-maker's beliefs, but also on the context, i.e. the current realization of the time series.

研究动机与目标

  • 建立一个决策模型,以应对时间压力,要求决策者在负面事件发生前做出预测。
  • 纳入内生信息获取机制,即决策者自主选择从随机时间序列中采样高成本信息的时机。
  • 刻画最优策略,以平衡信息增益(意外)与事件发生风险(悬念)的权衡。
  • 表明最优停止与继续决策不仅依赖于后验信念,还依赖于当前情境(时间序列的实现)。

提出的方法

  • 将决策问题形式化为一个带有随机截止时间(负面事件发生时间)的序贯贝叶斯推断任务,该截止时间会终止过程。
  • 将决策者的信念过程建模为一个下漂的超级 martingale,其因生存证据而随时间推移而下降。
  • 推导出具有‘会合’结构的最优策略,即在每次观测后可计算下一次采样时间。
  • 引入意外-悬念权衡:策略优化信息增益(通过信息增益的尾部分布)与生存概率(通过生存函数)的组合。
  • 使用观测成本($C_o$)、错误预测成本($C_1$)和终止风险成本($C_r$)来定义目标函数。
  • 将状态空间中的继续与停止区域表征为信念和当前时间序列实现($\bar{X}(P^\pi_t)$)的函数,表明其对情境的依赖性。

实验结果

研究问题

  • RQ1当决策过程因负面事件而终止时,决策者应如何最优地安排信息采样的时机?
  • RQ2在时间压力下,内生信息获取的最优策略具有何种结构?
  • RQ3意外(信息增益)与悬念(事件发生风险)在确定最优采样间隔时如何权衡?
  • RQ4继续与停止决策在多大程度上依赖于当前情境(即时间序列实现),而不仅仅是后验信念?
  • RQ5最优策略能否被表征为一种‘会合’策略,即在每次观测后可计算下一次采样时间?

主要发现

  • 最优策略具有‘会合’结构,意味着在每次观测后可立即计算出下一次信息采样的最优时间。
  • 最优采样间隔在意外(信息增益)与悬念(间隔内事件发生概率)之间实现平衡,该权衡由生存函数与信息增益分布共同捕捉。
  • 状态空间中的继续与停止区域同时依赖于后验信念和当前时间序列实现,显示出对情境的依赖性。
  • 当策略决定停止时,会基于阈值信念 $\frac{C_1}{C_o + C_1}$ 发出预测,该值权衡了错误预测的成本。
  • 在示例中,期望信息增益在 $\delta = 42$ 处达到最大,但最优会合时间可能更早(例如 $\delta^* < 42$),以维持合理的生存概率。
  • 最优会合时间点的生存概率受 $C_r$ 约束,策略在此处平衡了该风险与信息价值,体现于对 $S_t(\Delta t)$ 和 $I_t(\delta)$ 的成本加权优化。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。