Skip to main content
QUICK REVIEW

[Paper Review] Multi-Objective Reinforcement Learning for Infectious Disease Control with Application to COVID-19 Spread

Runzhe Wan, Xinyu Zhang|arXiv (Cornell University)|Sep 9, 2020
COVID-19 epidemiological studies45 references4 citations
TL;DR

This paper proposes a model-based multi-objective reinforcement learning framework that integrates a Bayesian epidemiological model with Pareto-optimal policy planning to balance disease control and economic costs during infectious disease outbreaks. Applied to China’s COVID-19 data, it enables real-time decision support by generating prediction bands for multiple intervention strategies, minimizing long-term societal costs.

ABSTRACT

Severe infectious diseases such as the novel coronavirus (COVID-19) pose a huge threat to public health. Stringent control measures, such as school closures and stay-at-home orders, while having significant effects, also bring huge economic losses. A crucial question for policymakers around the world is how to make the trade-off and implement the appropriate interventions. In this work, we propose a Multi-Objective Reinforcement Learning framework to facilitate the data-driven decision making and minimize the long-term overall cost. Specifically, at each decision point, a Bayesian epidemiological model is first learned as the environment model, and then we use the proposed model-based multi-objective planning algorithm to find a set of Pareto-optimal policies. This framework, combined with the prediction bands for each policy, provides a real-time decision support tool for policymakers. The application is demonstrated with the spread of COVID-19 in China.

Motivation & Objective

  • To address the challenge of balancing public health outcomes and economic impacts during infectious disease outbreaks.
  • To develop a data-driven decision support system that accounts for uncertainty in disease transmission and intervention effects.
  • To enable policymakers to evaluate trade-offs between health and economic costs using Pareto-optimal intervention strategies.
  • To provide real-time, prediction-aware policy recommendations through uncertainty quantification via prediction bands.

Proposed method

  • A Bayesian epidemiological model is learned at each decision point to represent the current state of disease transmission.
  • A model-based multi-objective planning algorithm is used to identify a set of Pareto-optimal policies that balance competing objectives.
  • The framework incorporates prediction bands for each policy to quantify uncertainty in outcomes over time.
  • The environment model is updated iteratively using observed data, enabling online adaptation to changing epidemic dynamics.
  • Multi-objective optimization is performed over long-term costs, including both infection burden and economic disruption.
  • The approach supports real-time decision making by ranking policies based on their trade-off performance across objectives.

Experimental results

Research questions

  • RQ1How can multi-objective reinforcement learning be used to identify intervention strategies that balance disease control and economic costs during an outbreak?
  • RQ2What role does uncertainty quantification via prediction bands play in supporting robust policy decisions under epidemic dynamics?
  • RQ3How effective is the proposed model-based framework in minimizing long-term societal costs compared to single-objective or heuristic approaches?
  • RQ4Can the framework generate actionable, Pareto-optimal policies that reflect real-world trade-offs in public health interventions?

Key findings

  • The framework successfully identifies a set of Pareto-optimal policies that balance infection control and economic impact during the COVID-19 outbreak in China.
  • Prediction bands for each policy provide reliable uncertainty quantification, enhancing trust and interpretability for policymakers.
  • The model-based approach enables real-time adaptation to evolving epidemic conditions through iterative environment model updates.
  • The integration of Bayesian modeling with multi-objective planning leads to more robust and data-driven policy recommendations than single-objective alternatives.
  • The method demonstrates practical utility by generating actionable intervention strategies with quantified trade-offs for decision-makers.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.