Skip to main content
QUICK REVIEW

[Paper Review] Partially Observable Risk-Sensitive Stopping Problems in Discrete Time

Nicole Bäuerle, Ulrich Rieder|arXiv (Cornell University)|Mar 28, 2017
Risk and Portfolio Optimization19 references3 citations
TL;DR

This paper develops a general framework for solving partially observable risk-sensitive optimal stopping problems in discrete time, using certainty equivalents derived from concave utility functions to model risk preferences. It proves the existence of optimal stopping times and shows that increased risk aversion leads to later stopping, with explicit recursive algorithms for computing reservation levels in the Bayesian house selling problem under exponential utility.

ABSTRACT

In this paper we consider stopping problems with partial observation under a general risk-sensitive optimization criterion for problems with finite and infinite time horizon. Our aim is to maximize the certainty equivalent of the stopping reward. We develop a general theory and discuss the Bayesian risk-sensitive house selling problem as a special example. In particular we are able to study the influence of the attitude towards risk of the decision maker on the optimal stopping rule.

Motivation & Objective

  • To extend risk-sensitive optimal stopping theory to partially observable Markov processes with finite and infinite time horizons.
  • To model decision-making under risk using the certainty equivalent of a random reward, where utility functions capture risk preferences.
  • To establish conditions for the existence of optimal stopping times and develop recursive algorithms for their computation.
  • To analyze how risk aversion influences optimal stopping rules, particularly in the Bayesian house selling problem.
  • To generalize existing risk-neutral results to the risk-sensitive case with partial information and exponential utility.

Proposed method

  • Uses the certainty equivalent $ \rho_U(X) = U^{-1}(\mathbb{E}[U(X)]) $ as the optimization criterion, where $ U $ is a concave utility function.
  • Applies the Arrow-Pratt measure of absolute risk aversion $ l_U(x) = -U''(x)/U'(x) $ to characterize risk sensitivity.
  • Develops a recursive value iteration algorithm for finite-horizon problems using the belief state $ \mu_n $ and accumulated cost $ s_n $.
  • Introduces an updating operator $ \Phi $ to propagate posterior beliefs after each observation in the partially observed system.
  • Derives fixed-point equations for reservation levels in the infinite-horizon case, particularly under exponential utility $ U(x) = \frac{1}{\gamma}e^{\gamma x} $ with $ \gamma < 0 $.
  • Proves optimality of stopping rules based on time- and belief-dependent reservation levels $ x_{n,\infty}^*(\mu) $, derived from backward induction.

Experimental results

Research questions

  • RQ1How can risk-sensitive optimal stopping be formulated in partially observable discrete-time systems with general utility functions?
  • RQ2What is the impact of risk aversion on the optimal stopping time in Bayesian models with incomplete information?
  • RQ3Can reservation levels be characterized in the infinite-horizon case, and how do they depend on the belief state and time?
  • RQ4How does the choice of utility function affect the structure of the optimal stopping rule?
  • RQ5Under what conditions does the optimal stopping time exist and can it be computed recursively?

Key findings

  • The optimal stopping time is characterized by time- and belief-dependent reservation levels $ x_{n,\infty}^*(\mu) $, which are computed via a recursive fixed-point equation.
  • For exponential utility $ U(x) = \frac{1}{\gamma}e^{\gamma x} $ with $ \gamma < 0 $, the reservation level satisfies $ x_{\infty}^*(\mu) = -c + \frac{1}{\gamma}\ln\left(\int e^{\gamma \max\{x, x_{\infty}^*(\Phi(x,\mu))\}} Q(dx|\mu)\right) $.
  • The optimal stopping time is almost surely finite under general conditions, even with strictly negative running costs.
  • More risk-averse decision makers (lower $ \gamma $) stop later, as shown by the monotonicity of $ x_{n,\infty}^*(\mu, U) $ in risk aversion.
  • Reservation levels $ x_{n,\infty}^*(\mu) $ are decreasing in time $ n $ when the utility function is DARA (decreasing absolute risk aversion).
  • The value function is bounded and satisfies a dynamic programming principle, enabling recursive computation of the optimal policy.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.