[Paper Review] Multi-Objective Reinforcement Learning for Infectious Disease Control with Application to COVID-19 Spread
This paper proposes a model-based multi-objective reinforcement learning framework that integrates a Bayesian epidemiological model with Pareto-optimal policy planning to balance disease control and economic costs during infectious disease outbreaks. Applied to China’s COVID-19 data, it enables real-time decision support by generating prediction bands for multiple intervention strategies, minimizing long-term societal costs.
Severe infectious diseases such as the novel coronavirus (COVID-19) pose a huge threat to public health. Stringent control measures, such as school closures and stay-at-home orders, while having significant effects, also bring huge economic losses. A crucial question for policymakers around the world is how to make the trade-off and implement the appropriate interventions. In this work, we propose a Multi-Objective Reinforcement Learning framework to facilitate the data-driven decision making and minimize the long-term overall cost. Specifically, at each decision point, a Bayesian epidemiological model is first learned as the environment model, and then we use the proposed model-based multi-objective planning algorithm to find a set of Pareto-optimal policies. This framework, combined with the prediction bands for each policy, provides a real-time decision support tool for policymakers. The application is demonstrated with the spread of COVID-19 in China.
Motivation & Objective
- To address the challenge of balancing public health outcomes and economic impacts during infectious disease outbreaks.
- To develop a data-driven decision support system that accounts for uncertainty in disease transmission and intervention effects.
- To enable policymakers to evaluate trade-offs between health and economic costs using Pareto-optimal intervention strategies.
- To provide real-time, prediction-aware policy recommendations through uncertainty quantification via prediction bands.
Proposed method
- A Bayesian epidemiological model is learned at each decision point to represent the current state of disease transmission.
- A model-based multi-objective planning algorithm is used to identify a set of Pareto-optimal policies that balance competing objectives.
- The framework incorporates prediction bands for each policy to quantify uncertainty in outcomes over time.
- The environment model is updated iteratively using observed data, enabling online adaptation to changing epidemic dynamics.
- Multi-objective optimization is performed over long-term costs, including both infection burden and economic disruption.
- The approach supports real-time decision making by ranking policies based on their trade-off performance across objectives.
Experimental results
Research questions
- RQ1How can multi-objective reinforcement learning be used to identify intervention strategies that balance disease control and economic costs during an outbreak?
- RQ2What role does uncertainty quantification via prediction bands play in supporting robust policy decisions under epidemic dynamics?
- RQ3How effective is the proposed model-based framework in minimizing long-term societal costs compared to single-objective or heuristic approaches?
- RQ4Can the framework generate actionable, Pareto-optimal policies that reflect real-world trade-offs in public health interventions?
Key findings
- The framework successfully identifies a set of Pareto-optimal policies that balance infection control and economic impact during the COVID-19 outbreak in China.
- Prediction bands for each policy provide reliable uncertainty quantification, enhancing trust and interpretability for policymakers.
- The model-based approach enables real-time adaptation to evolving epidemic conditions through iterative environment model updates.
- The integration of Bayesian modeling with multi-objective planning leads to more robust and data-driven policy recommendations than single-objective alternatives.
- The method demonstrates practical utility by generating actionable intervention strategies with quantified trade-offs for decision-makers.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.