[Paper Review] Entropy-regularized penalization schemes and reflected BSDEs with singular generators
Introduces an entropy-regularized penalization scheme for continuous-time optimal stopping, analyzes convergence to American options, and derives a limit RBSDE with a logarithmically singular driver.
This paper extends our previous work to continuous-time optimal stopping, focusing on American options in an exploratory setting. Our first contribution is an entropy-regularized penalization scheme, inspired by classical penalization techniques for reflected BSDEs. It yields a smooth approximation of the stopping rule, promotes exploration, and enables gradient-based learning methods. We prove well-posedness, convergence, and illustrate numerical performance in low-dimensional examples. Our second contribution analyzes the behaviour of the scheme as the penalization parameter grows, showing that the limit solves a reflected BSDE with a logarithmically singular generator, for which we establish existence and uniqueness via a monotone limit argument.
Motivation & Objective
- Extend entropy-regularized penalization to continuous-time optimal stopping and American options in an exploratory setting.
- Establish well-posedness and convergence of the regularized scheme as penalization and temperature parameters vary.
- Develop a Policy Improvement Algorithm (PIA) tailored to the regularized framework.
- Analyze the limiting behavior as the penalization parameter grows to infinity and identify the resulting singular RBSDE.
- Provide numerical illustrations and connect the scheme to probabilistic interpretations of optimal stopping.
Proposed method
- Formulate an entropy-regularized BSDE to approximate the Snell envelope under a relaxed control framework.
- Use a Gibbs-type representation for the optimal control and derive a Lipschitz driver for fixed truncation n and temperature λ.
- Prove well-posedness of the entropy-regularized BSDE and analyze convergence to the classical penalized value V^n as λ→0.
- Show that V^{λ,n} converges to the American option value V when λ ln(n) → 0.
- Define a Policy Improvement Algorithm (PIA) to iteratively improve the policy and evaluate the value function within the regularized setup.
- Investigate the limit as n→∞ to obtain a reflected BSDE with a logarithmically singular driver and establish existence/uniqueness via monotone limit arguments.
Experimental results
Research questions
- RQ1How does entropy regularization affect the regularity and tractability of continuous-time optimal stopping in an exploratory setting?
- RQ2Do entropy-regularized schemes converge to the classical American option value when penalization and temperature parameters are tuned appropriately?
- RQ3What is the limiting behavior of the scheme as the penalization parameter grows without bound, and can we characterize the limit as a singular RBSDE?
- RQ4Can a Policy Improvement Algorithm be effectively integrated into the entropy-regularized framework for practical computation?
- RQ5What probabilistic interpretation emerges for the limit process in terms of exploration and risk considerations?
Key findings
- The entropy-regularized scheme yields a well-posed BSDE with a Lipschitz driver for fixed n and λ.
- As λ→0 with λ ln(n)→0, V^{λ,n} converges to the classical penalized value V^n and further to the American option value V under suitable scaling.
- A monotone limit as n→∞ leads to a reflected BSDE with a logarithmically singular driver, and existence/uniqueness hold via monotone arguments.
- The paper provides a Policy Improvement Algorithm that generates a monotonically increasing sequence of value estimates V^{λ,m} under the regularized objective.
- Numerical experiments illustrate the practical performance of the regularized scheme and the PIA in low-dimensional settings.
- The limiting singular driver is analyzed, and the results give a probabilistic interpretation linking entropy regularization to endogenous default risk and early exercise.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.