Skip to main content
QUICK REVIEW

[论文解读] Reinforcement Learning for Portfolio Management

Angelos Filos|arXiv (Cornell University)|Sep 12, 2019
Stock Market Forecasting Methods参考文献 64被引用 9
一句话总结

本文提出无需模型的强化学习智能体——具体为深度软循环Q网络(DSRQN)和得分机混合模型(MSM)——用于跨多种金融资产和市场的通用投资组合管理。通过利用数据增强和预训练,这些智能体在交易成本和非平稳市场条件下仍表现出色,相较于最先进模型,年化累计收益最高提升9.2%,年化夏普比率最高提升13.4%。

ABSTRACT

In this thesis, we develop a comprehensive account of the expressive power, modelling efficiency, and performance advantages of so-called trading agents (i.e., Deep Soft Recurrent Q-Network (DSRQN) and Mixture of Score Machines (MSM)), based on both traditional system identification (model-based approach) as well as on context-independent agents (model-free approach). The analysis provides conclusive support for the ability of model-free reinforcement learning methods to act as universal trading agents, which are not only capable of reducing the computational and memory complexity (owing to their linear scaling with the size of the universe), but also serve as generalizing strategies across assets and markets, regardless of the trading universe on which they have been trained. The relatively low volume of daily returns in financial market data is addressed via data augmentation (a generative approach) and a choice of pre-training strategies, both of which are validated against current state-of-the-art models. For rigour, a risk-sensitive framework which includes transaction costs is considered, and its performance advantages are demonstrated in a variety of scenarios, from synthetic time-series (sinusoidal, sawtooth and chirp waves), simulated market series (surrogate data based), through to real market data (S\&P 500 and EURO STOXX 50). The analysis and simulations confirm the superiority of universal model-free reinforcement learning agents over current portfolio management model in asset allocation strategies, with the achieved performance advantage of as much as 9.2\% in annualized cumulative returns and 13.4\% in annualized Sharpe Ratio.

研究动机与目标

  • 开发可跨不同金融资产和市场泛化、无需微调的通用无模型强化学习智能体。
  • 通过数据增强和预训练应对金融时间序列中的非平稳性、弱历史关联性及数据量不足的挑战。
  • 在包含交易成本和多样化市场状态的风险敏感环境中评估这些智能体的性能。
  • 在合成数据、代理数据和真实市场数据上,将所提智能体与最先进投资组合优化模型进行对比。

提出的方法

  • 采用深度软循环Q网络(DSRQN)和得分机混合模型(MSM)作为无模型强化学习智能体,用于序列资产配置。
  • 通过生成式方法实施数据增强,以增加有效训练数据量,并提升低数据场景下的泛化能力。
  • 实施预训练策略,以提升在稀疏收益金融时间序列中的学习效率与收敛性。
  • 在强化学习目标函数中显式整合交易成本的风险敏感框架。
  • 在多样化金融标的池(标普500、欧元区斯托克50)上训练智能体,并评估其在未见资产与市场中的泛化能力。
  • 使用合成时间序列(正弦、锯齿、啁啾波)和代理数据,验证在受控与非平稳条件下的鲁棒性。

实验结果

研究问题

  • RQ1无模型强化学习智能体是否能在不重新训练的情况下,跨不同金融资产和市场实现泛化?
  • RQ2数据增强与预训练在低样本金融收益数据中能在多大程度上提升性能?
  • RQ3与最先进模型相比,所提智能体在包含交易成本等风险敏感约束下的表现如何?
  • RQ4与传统基于模型和无模型的投资组合优化方法相比,通用强化学习智能体的性能优势是什么?
  • RQ5在合成数据、代理数据和真实市场数据中,智能体在累计收益与夏普比率方面的泛化能力如何?

主要发现

  • 所提无模型强化学习智能体相较于最先进投资组合管理模型,年化累计收益最高提升9.2%。
  • 智能体在年化夏普比率上实现13.4%的提升,表明其风险调整后表现更优。
  • 智能体在无需微调的情况下,能有效泛化至不同金融标的池,包括标普500与欧元区斯托克50。
  • 在风险敏感框架中整合交易成本,显著提升了模型的鲁棒性与实际应用价值。
  • 数据增强与预训练策略有效缓解了金融时间序列中的数据稀缺问题,增强了学习稳定性与性能表现。
  • 在合成数据、代理数据与真实市场数据中,性能优势均持续显现,证实了该方法的鲁棒性与通用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。