Skip to main content
QUICK REVIEW

[Paper Review] A Survey on Practical Applications of Multi-Armed and Contextual Bandits

Djallel Bouneffouf, Irina Rish|arXiv (Cornell University)|Apr 2, 2019
Advanced Bandit Algorithms Research20 references106 citations
TL;DR

This survey reviews practical applications of multi-armed and contextual bandits across healthcare, finance, pricing, recommender systems, and more, and discusses how bandit methods inform real-world decision making and machine learning workflows.

ABSTRACT

In recent years, multi-armed bandit (MAB) framework has attracted a lot of attention in various applications, from recommender systems and information retrieval to healthcare and finance, due to its stellar performance combined with certain attractive properties, such as learning from less feedback. The multi-armed bandit field is currently flourishing, as novel problem settings and algorithms motivated by various practical applications are being introduced, building on top of the classical bandit problem. This article aims to provide a comprehensive review of top recent developments in multiple real-life applications of the multi-armed bandit. Specifically, we introduce a taxonomy of common MAB-based applications and summarize state-of-art for each of those domains. Furthermore, we identify important current trends and provide new perspectives pertaining to the future of this exciting and fast-growing field.

Motivation & Objective

  • Provide a taxonomy of real-world MAB and contextual bandit applications across domains.
  • Summarize state-of-the-art algorithms used in each domain and their advantages.
  • Identify trends, gaps, and open problems to guide future research in bandit methods.

Proposed method

  • Describe the standard MAB and contextual bandit frameworks and their relevance to real-world settings.
  • Review domain-specific applications and the corresponding bandit formulations used (MAB vs CMAB, stationary vs non-stationary).
  • Highlight notable algorithms and modeling approaches (e.g., LINUCB, CTS, Thompson Sampling, bandit with side information).
  • Discuss how bandits can augment machine learning workflows, including hyperparameter tuning, feature selection, active learning, and RL orchestration.

Experimental results

Research questions

  • RQ1What are the main real-world domains where MAB and CMAB have been effectively applied?
  • RQ2Which bandit formulations and algorithms are most successful in each domain?
  • RQ3What are the identified gaps and opportunities for future bandit research and cross-domain transfer?
  • RQ4How can bandit methods enhance broader machine learning tasks such as hyperparameter optimization and active learning?

Key findings

  • A broad taxonomy of practical MAB and CMAB applications spans healthcare, finance, dynamic pricing, recommender systems, influence maximization, information retrieval, dialogue systems, anomaly detection, and telecommunications.
  • Contextual bandits and non-stationary variants are used in several domains, with specific choices like LINUCB, CTS, and Thompson Sampling guiding decisions.
  • Bandits provide advantages in online decision making with limited feedback and exploration needs, informing adaptive experimentation and learning in real time.
  • There is limited cross-domain transfer or multitask bandit work, suggesting opportunities for lifelong learning and transfer learning in bandit settings.
  • Bandits can augment machine learning pipelines, including algorithm selection, hyperparameter optimization (e.g., Hyperband), feature selection, active learning, clustering, and online RL orchestration.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.