Skip to main content
QUICK REVIEW

[Paper Review] Entropy-regularized penalization schemes and reflected BSDEs with singular generators

Daniel Chee, Noufel Frikha|arXiv (Cornell University)|Feb 20, 2026
Stochastic processes and financial applications0 citations
TL;DR

Introduces an entropy-regularized penalization scheme for continuous-time optimal stopping, analyzes convergence to American options, and derives a limit RBSDE with a logarithmically singular driver.

ABSTRACT

This paper extends our previous work to continuous-time optimal stopping, focusing on American options in an exploratory setting. Our first contribution is an entropy-regularized penalization scheme, inspired by classical penalization techniques for reflected BSDEs. It yields a smooth approximation of the stopping rule, promotes exploration, and enables gradient-based learning methods. We prove well-posedness, convergence, and illustrate numerical performance in low-dimensional examples. Our second contribution analyzes the behaviour of the scheme as the penalization parameter grows, showing that the limit solves a reflected BSDE with a logarithmically singular generator, for which we establish existence and uniqueness via a monotone limit argument.

Motivation & Objective

  • Extend entropy-regularized penalization to continuous-time optimal stopping and American options in an exploratory setting.
  • Establish well-posedness and convergence of the regularized scheme as penalization and temperature parameters vary.
  • Develop a Policy Improvement Algorithm (PIA) tailored to the regularized framework.
  • Analyze the limiting behavior as the penalization parameter grows to infinity and identify the resulting singular RBSDE.
  • Provide numerical illustrations and connect the scheme to probabilistic interpretations of optimal stopping.

Proposed method

  • Formulate an entropy-regularized BSDE to approximate the Snell envelope under a relaxed control framework.
  • Use a Gibbs-type representation for the optimal control and derive a Lipschitz driver for fixed truncation n and temperature λ.
  • Prove well-posedness of the entropy-regularized BSDE and analyze convergence to the classical penalized value V^n as λ→0.
  • Show that V^{λ,n} converges to the American option value V when λ ln(n) → 0.
  • Define a Policy Improvement Algorithm (PIA) to iteratively improve the policy and evaluate the value function within the regularized setup.
  • Investigate the limit as n→∞ to obtain a reflected BSDE with a logarithmically singular driver and establish existence/uniqueness via monotone limit arguments.

Experimental results

Research questions

  • RQ1How does entropy regularization affect the regularity and tractability of continuous-time optimal stopping in an exploratory setting?
  • RQ2Do entropy-regularized schemes converge to the classical American option value when penalization and temperature parameters are tuned appropriately?
  • RQ3What is the limiting behavior of the scheme as the penalization parameter grows without bound, and can we characterize the limit as a singular RBSDE?
  • RQ4Can a Policy Improvement Algorithm be effectively integrated into the entropy-regularized framework for practical computation?
  • RQ5What probabilistic interpretation emerges for the limit process in terms of exploration and risk considerations?

Key findings

  • The entropy-regularized scheme yields a well-posed BSDE with a Lipschitz driver for fixed n and λ.
  • As λ→0 with λ ln(n)→0, V^{λ,n} converges to the classical penalized value V^n and further to the American option value V under suitable scaling.
  • A monotone limit as n→∞ leads to a reflected BSDE with a logarithmically singular driver, and existence/uniqueness hold via monotone arguments.
  • The paper provides a Policy Improvement Algorithm that generates a monotonically increasing sequence of value estimates V^{λ,m} under the regularized objective.
  • Numerical experiments illustrate the practical performance of the regularized scheme and the PIA in low-dimensional settings.
  • The limiting singular driver is analyzed, and the results give a probabilistic interpretation linking entropy regularization to endogenous default risk and early exercise.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.