Skip to main content
QUICK REVIEW

[Paper Review] Adaptive Combinatorial Experimental Design: Pareto Optimality for Decision-Making and Inference

Hongrui Xie, Junyu Cao|arXiv (Cornell University)|Feb 27, 2026
Advanced Bandit Algorithms Research0 citations
TL;DR

This paper introduces Pareto-optimal learning for adaptive combinatorial bandits, proposes MixCombKL (full-bandit) and MixCombUCB (semi-bandit), and proves finite-time regret and gap-estimation guarantees under two feedback regimes.

ABSTRACT

In this paper, we provide the first investigation into adaptive combinatorial experimental design, focusing on the trade-off between regret minimization and statistical power in combinatorial multi-armed bandits (CMAB). While minimizing regret requires repeated exploitation of high-reward arms, accurate inference on reward gaps requires sufficient exploration of suboptimal actions. We formalize this trade-off through the concept of Pareto optimality and establish equivalent conditions for Pareto-efficient learning in CMAB. We consider two relevant cases under different information structures, i.e., full-bandit feedback and semi-bandit feedback, and propose two algorithms MixCombKL and MixCombUCB respectively for these two cases. We provide theoretical guarantees showing that both algorithms are Pareto optimal, achieving finite-time guarantees on both regret and estimation error of arm gaps. Our results further reveal that richer feedback significantly tightens the attainable Pareto frontier, with the primary gains arising from improved estimation accuracy under our proposed methods. Taken together, these findings establish a principled framework for adaptive combinatorial experimentation in multi-objective decision-making.

Motivation & Objective

  • Motivate the study of regret versus inference trade-offs in combinatorial bandits (CMAB).
  • Formalize Pareto optimality as a framework for balancing regret and reward-gap estimation.
  • Develop Pareto-optimal algorithms for two feedback models (full-bandit and semi-bandit).
  • Provide finite-time guarantees for both regret and estimation errors under each feedback regime.

Proposed method

  • Model CMAB with base arms and super arms under full-bandit and semi-bandit feedback.
  • Define Pareto optimality and Pareto frontier to capture trade-offs between regret and estimation error.
  • Develop MixCombKL for full-bandit feedback using KL-divergence guided online stochastic mirror descent on a simplex embedding.
  • Develop MixCombUCB for semi-bandit feedback using a UCB-based approach with an initialization phase and an optimization oracle.
  • Provide finite-sample bounds for estimation errors of super-arm gaps and base-arm gaps, and regret bounds for both algorithms.
  • Establish necessary and sufficient conditions for Pareto optimality and relate richness of feedback to frontier tightness.

Experimental results

Research questions

  • RQ1What is the trade-off between regret minimization and statistical power for estimating reward gaps in CMAB?
  • RQ2Can Pareto optimal policies be characterized and achieved in CMAB under different feedback regimes?
  • RQ3How do full-bandit and semi-bandit feedback impact the Pareto frontier and learning guarantees?
  • RQ4What are the finite-time estimation and regret guarantees for MixCombKL and MixCombUCB?
  • RQ5What are the necessary and sufficient conditions for Pareto optimality in CMAB settings?

Key findings

  • MixCombKL achieves Pareto-optimal trade-offs under full-bandit feedback with finite-time gap estimation guarantees and regret bounds.
  • MixCombUCB achieves Pareto-optimal trade-offs under semi-bandit feedback with finite-time gap estimation guarantees and regret bounds.
  • Semi-bandit feedback yields a sharper Pareto frontier than full-bandit feedback due to improved estimation accuracy, while regret scales similarly under the proposed algorithms.
  • The paper provides explicit finite-sample bounds for both super-arm gap estimation and base-arm gap estimation, along with regret bounds.
  • Pareto optimality is characterized by conditions linking estimation error and regret, applicable to both feedback models.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.