[Paper Review] Least Squares Policy Iteration with Instrumental Variables vs. Direct Policy Search: Comparison Against Optimal Benchmarks Using Energy Storage
This paper compares least-squares policy iteration with instrumental variables (IV) against direct policy search in energy storage applications, demonstrating that IV-based least-squares policy iteration significantly outperforms standard least-squares Bellman error minimization but still underperforms direct policy search using knowledge gradient. The study establishes theoretical equivalence between IV-based Bellman error minimization and projected Bellman error minimization, providing a consistent and efficient alternative to conventional approximate dynamic programming methods in high-dimensional stochastic control problems.
This paper studies approximate policy iteration (API) methods which use least-squares Bellman error minimization for policy evaluation. We address several of its enhancements, namely, Bellman error minimization using instrumental variables, least-squares projected Bellman error minimization, and projected Bellman error minimization using instrumental variables. We prove that for a general discrete-time stochastic control problem, Bellman error minimization using instrumental variables is equivalent to both variants of projected Bellman error minimization. An alternative to these API methods is direct policy search based on knowledge gradient. The practical performance of these three approximate dynamic programming methods are then investigated in the context of an application in energy storage, integrated with an intermittent wind energy supply to fully serve a stochastic time-varying electricity demand. We create a library of test problems using real-world data and apply value iteration to find their optimal policies. These benchmarks are then used to compare the developed policies. Our analysis indicates that API with instrumental variables Bellman error minimization prominently outperforms API with least-squares Bellman error minimization. However, these approaches underperform our direct policy search implementation.
Motivation & Objective
- To evaluate the performance of approximate policy iteration (API) methods using least-squares Bellman error minimization with instrumental variables (IV) in energy storage applications.
- To compare IV-based LSPI against standard least-squares policy iteration and direct policy search using knowledge gradient in a realistic, high-dimensional stochastic control problem.
- To establish theoretical equivalence between instrumental variable-based Bellman error minimization and projected Bellman error minimization (MSPBE) in discrete-time stochastic control.
- To provide optimal benchmarks via value iteration on real-world energy storage datasets to rigorously assess algorithmic performance.
- To investigate the practical scalability and accuracy of different approximate dynamic programming strategies in the context of renewable energy integration.
Proposed method
- Uses least-squares policy iteration (LSPI) with instrumental variables (IV) to minimize Bellman error, ensuring consistent estimation even when regressors are endogenous.
- Proves mathematically that IV-based Bellman error minimization is equivalent to mean-squared projected Bellman error (MSPBE) minimization under general conditions.
- Applies linear value function approximation to model energy storage policies, using real-world wind and load data to generate benchmark problems.
- Employs value iteration on the generated datasets to compute optimal policies, serving as ground-truth benchmarks for algorithmic comparison.
- Implements direct policy search via knowledge gradient to compare against API-based methods in terms of policy quality and convergence.
- Uses statistical consistency theory and asymptotic analysis to validate the convergence of IV estimators under standard assumptions, including exogeneity and full rank of instrumental variable covariance.
Experimental results
Research questions
- RQ1Is instrumental variable-based least-squares policy iteration more consistent and accurate than standard least-squares Bellman error minimization in high-dimensional energy storage problems?
- RQ2Does the theoretical equivalence between IV-based Bellman error minimization and MSPBE minimization hold in practice for stochastic control problems with endogenous regressors?
- RQ3How does direct policy search using knowledge gradient compare in performance to approximate policy iteration with IV-based or standard LSPI methods?
- RQ4Can optimal benchmarks be reliably generated for energy storage problems using value iteration on real-world data, enabling rigorous evaluation of approximate dynamic programming methods?
- RQ5What is the relative performance gap between LSPI with IV and direct policy search when applied to a realistic, time-varying electricity demand and intermittent wind supply scenario?
Key findings
- Instrumental variable-based least-squares policy iteration (IV-LSPI) significantly outperforms standard least-squares policy iteration (LSPI) in terms of policy quality, as measured against optimal benchmarks.
- Theoretical analysis proves that IV-based Bellman error minimization is mathematically equivalent to mean-squared projected Bellman error (MSPBE) minimization, validating the use of IV as a consistent estimator in approximate dynamic programming.
- Despite its theoretical advantages, IV-LSPI still underperforms direct policy search using knowledge gradient, indicating that direct policy search may be more effective for this class of problems.
- The study generates a library of real-world energy storage test problems using actual wind and load data, enabling reproducible and rigorous evaluation of algorithmic performance.
- Value iteration is successfully applied to solve the benchmark problems, providing optimal policies that serve as a gold standard for evaluating approximate methods.
- The results highlight a performance gap between approximate policy iteration and direct policy search, suggesting that direct search strategies may be more suitable for complex, high-dimensional energy storage control problems.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.