Skip to main content
QUICK REVIEW

[Paper Review] Deep Reinforcement Learning for Optimum Order Execution: Mitigating Risk and Maximizing Returns

Khabbab Zakaria, Jayapaulraj Jerinsh|arXiv (Cornell University)|Jan 8, 2026
Risk and Portfolio Optimization0 citations
TL;DR

The paper presents a Deep Reinforcement Learning approach for optimal order execution in the US market, outperforming VWAP and TWAP in ROI and risk management by dynamically adapting to market conditions, including stress periods.

ABSTRACT

Optimal Order Execution is a well-established problem in finance that pertains to the flawless execution of a trade (buy or sell) for a given volume within a specified time frame. This problem revolves around optimizing returns while minimizing risk, yet recent research predominantly focuses on addressing one aspect of this challenge. In this paper, we introduce an innovative approach to Optimal Order Execution within the US market, leveraging Deep Reinforcement Learning (DRL) to effectively address this optimization problem holistically. Our study assesses the performance of our model in comparison to two widely employed execution strategies: Volume Weighted Average Price (VWAP) and Time Weighted Average Price (TWAP). Our experimental findings clearly demonstrate that our DRL-based approach outperforms both VWAP and TWAP in terms of return on investment and risk management. The model's ability to adapt dynamically to market conditions, even during periods of market stress, underscores its promise as a robust solution.

Motivation & Objective

  • Motivate the need to optimize returns while minimizing risk in optimal order execution.
  • Propose a holistic DRL-based framework for US market execution.
  • Evaluate DRL against VWAP and TWAP across performance and risk metrics.
  • Demonstrate robustness of DRL during stressed market conditions.

Proposed method

  • Develop a Deep Reinforcement Learning model for optimal order execution in the US market.
  • Compare DRL performance against VWAP and TWAP execution strategies.
  • Assess outcomes in terms of return on investment and risk management.
  • Test adaptability of the model under market stress and varying conditions.
  • Focus on dynamic adaptation to evolving market environments.

Experimental results

Research questions

  • RQ1Can a DRL-based optimizer outperform VWAP and TWAP in ROI and risk metrics for US market executions?
  • RQ2How does the DRL model adapt to changing market conditions and during market stress?
  • RQ3What are the key factors driving performance differences between DRL and traditional execution strategies?

Key findings

  • DRL-based optimum order execution outperforms VWAP and TWAP in ROI and risk management.
  • The DRL model demonstrates dynamic adaptation to market conditions.
  • The approach remains robust during periods of market stress.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.