Skip to main content
QUICK REVIEW

[Paper Review] Estimation of Optimal Dynamic Treatment Assignment Rules under Policy Constraints

Shosei Sakaguchi|arXiv (Cornell University)|Jun 9, 2021
Advanced Causal Inference Techniques4 citations
TL;DR

This paper proposes a Dynamic Empirical Welfare Maximization (DEWM) framework for estimating optimal dynamic treatment rules under policy constraints using panel data from experimental or quasi-experimental studies. It introduces two estimation methods—backward induction and simultaneous optimization—achieving $n^{-1/2}$-minimax convergence rates and finite-sample welfare-regret bounds, with extensions for intertemporal budget and capacity constraints.

ABSTRACT

Many policies involve dynamics in their treatment assignments, where individuals receive sequential interventions over multiple stages. We study estimation of an optimal dynamic treatment regime that guides the optimal treatment assignment for each individual at each stage based on their history. We propose an empirical welfare maximization approach in this dynamic framework, which estimates the optimal dynamic treatment regime using data from an experimental or quasi-experimental study while satisfying exogenous constraints on policies. The paper proposes two estimation methods: one solves the treatment assignment problem sequentially through backward induction, and the other solves the entire problem simultaneously across all stages. We establish finite-sample upper bounds on worst-case average welfare regrets for these methods and show their optimal $n^{-1/2}$ convergence rates. We also modify the simultaneous estimation method to accommodate intertemporal budget/capacity constraints.

Motivation & Objective

  • To develop a statistical framework for estimating optimal dynamic treatment assignment rules in sequential decision-making contexts.
  • To address the challenge of heterogeneous treatment effects that depend on past treatments, outcomes, and covariates in dynamic settings.
  • To incorporate exogenous policy constraints—such as interpretability, fairness, or budget limits—into the estimation of dynamic treatment regimes.
  • To establish finite-sample welfare-regret bounds and minimax convergence rates for the proposed estimation methods.
  • To extend the empirical welfare maximization (EWM) approach to dynamic, multi-stage treatment problems with panel data.

Proposed method

  • Employs a Dynamic Empirical Welfare Maximization (DEWM) approach that maximizes empirical welfare over a pre-specified class of feasible dynamic treatment rules (DTRs).
  • Proposes two estimation strategies: backward induction, which solves the treatment choice problem stage-by-stage from the final stage to the first, and simultaneous optimization, which solves the full DTR estimation problem across all stages at once.
  • Uses propensity score-weighted outcomes in the empirical welfare objective to account for treatment assignment mechanisms in experimental and observational data.
  • Represents the DTR estimation problem as a mixed-integer linear program (MILP), enabling computation via standard optimization solvers under linear eligibility rule constraints.
  • Incorporates policy constraints—such as start/stop timing or one-shot treatment rules—by adding linear constraints to the MILP formulation.
  • Extends the simultaneous DEWM method to include intertemporal budget or capacity constraints through additional linear constraints in the MILP, ensuring resource limits are respected.

Experimental results

Research questions

  • RQ1How can optimal dynamic treatment rules be estimated in settings with sequential, heterogeneous treatment effects and observed covariates?
  • RQ2What are the finite-sample performance guarantees—specifically, welfare-regret bounds—for dynamic treatment rule estimation under policy constraints?
  • RQ3How do backward induction and simultaneous optimization methods compare in terms of statistical efficiency and computational feasibility?
  • RQ4Can the DEWM framework be adapted to incorporate real-world policy constraints such as budget limits or treatment duration restrictions?
  • RQ5What is the minimax convergence rate of the proposed estimators under regularity conditions?

Key findings

  • The proposed DEWM methods achieve $n^{-1/2}$-minimax convergence rates, indicating optimal statistical efficiency under standard regularity conditions.
  • Finite-sample upper bounds on worst-case average welfare regret are derived, providing theoretical guarantees on the performance of the estimated DTRs.
  • The backward induction method provides a recursive solution path that accounts for future treatment effects, improving long-term welfare optimization.
  • The simultaneous estimation method enables joint optimization across all stages, capturing complex interactions between treatment decisions at different time points.
  • The MILP formulation allows for efficient computation of optimal DTRs under various policy constraints, including start/stop timing and one-shot treatment rules.
  • The inclusion of budget or capacity constraints in the simultaneous DEWM framework is achieved via additional linear constraints in the MILP, preserving computational tractability.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.