Skip to main content
QUICK REVIEW

[Paper Review] An Explicit Example Of Optimal Portfolio-Consumption Choices With Habit Formation And Partial Observations

Xiang Yu|arXiv (Cornell University)|Dec 13, 2011
Stochastic processes and financial applications18 references3 citations
TL;DR

This paper provides an explicit solution to the optimal portfolio-consumption problem under habit formation and partial observations using Kalman-Bucy filtering and dynamic programming. It derives closed-form feedback strategies for investment and consumption in a diffusion market with unobservable drift, showing that partial information simplifies the path-dependent control problem compared to full information.

ABSTRACT

We consider a model of optimal investment and consumption with both habit formation and partial observations in incomplete Itô processes market. The investor chooses his consumption under the addictive habits constraint while only observing the market stock prices but not the instantaneous rate of return. Applying the Kalman-Bucy filtering theorem and the Dynamic Programming arguments, we solve the associated Hamilton-Jacobi-Bellman (HJB) equation explicitly for the path dependent stochastic control problem in the case of power utilities. We provide the optimal investment and consumption policies in explicit feedback forms using rigorous verification arguments.

Motivation & Objective

  • To model optimal investment and consumption under addictive habit formation, where current utility depends on past consumption levels.
  • To address the realistic constraint that investors only observe stock prices, not the unobservable instantaneous drift of the risky asset.
  • To solve the resulting path-dependent stochastic control problem under partial information explicitly, avoiding reliance on the Dynamic Programming Principle.
  • To demonstrate that partial observation can simplify the solution structure compared to full information settings.
  • To provide rigorous verification of explicit feedback forms for optimal policies using a verification theorem.

Proposed method

  • Model the stock price and unobservable drift as correlated Itô processes, with the drift following an Ornstein-Uhlenbeck process.
  • Apply the Kalman-Bucy filtering theorem to estimate the unobserved drift based on observed stock prices.
  • Formulate the Hamilton-Jacobi-Bellman (HJB) equation for the value function under power utility and partial information.
  • Decouple the HJB equation into auxiliary ordinary differential equations (ODEs) with constant coefficients via ansatz methods.
  • Use a verification theorem to rigorously confirm that the candidate feedback controls are optimal.
  • Solve the resulting system of ODEs under four distinct parameter regimes: hyperbolic, polynomial, tangent, and critical cases.

Experimental results

Research questions

  • RQ1How does habit formation affect optimal consumption and investment when the investor cannot observe the true drift of the risky asset?
  • RQ2Can explicit feedback forms for optimal policies be derived in a partial information setting with habit formation?
  • RQ3Why does the partial observation setting yield a simpler solution than the full information case in this path-dependent control problem?
  • RQ4What are the structural differences in optimal strategies across different parameter regimes (e.g., hyperbolic vs. polynomial vs. tangent solutions)?
  • RQ5Can the verification theorem be applied directly to confirm optimality without proving the Dynamic Programming Principle?

Key findings

  • The optimal consumption policy is explicitly expressed as a feedback function of the filtered drift and the current habit level, with a time-varying adjustment based on the investment horizon.
  • The optimal portfolio weight in the risky asset is derived in closed form as a function of the filtered drift and the time-to-maturity, using a linear feedback structure.
  • Four distinct solution types emerge—hyperbolic, polynomial, tangent, and critical—depending on the parameter configuration, each with different qualitative behaviors (bounded or explosive).
  • The hyperbolic solution is bounded when γ₂ < 0, while the polynomial solution is bounded under specific conditions on λ, p, ρ, σ_S, and σ_μ.
  • The tangent solution is explosive and exists when the discriminant Δ < 0, with a critical time point at which the solution diverges.
  • The verification theorem confirms the optimality of the derived policies without requiring the full proof of the Dynamic Programming Principle.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.