Skip to main content
QUICK REVIEW

[Paper Review] Fast Online Learning with Gaussian Prior-Driven Hierarchical Unimodal Thompson Sampling

Tianchi Zhao, He Liu|arXiv (Cornell University)|Feb 17, 2026
Advanced Bandit Algorithms Research0 citations
TL;DR

The paper introduces two Thompson Sampling variants (TSCG and UTSCG) for Gaussian-arm bandits with clustered and unimodal structure, proving tighter regret bounds and validating performance in mmWave and portfolio scenarios.

ABSTRACT

We study a type of Multi-Armed Bandit (MAB) problems in which arms with a Gaussian reward feedback are clustered. Such an arm setting finds applications in many real-world problems, for example, mmWave communications and portfolio management with risky assets, as a result of the universality of the Gaussian distribution. Based on the Thompson Sampling algorithm with Gaussian prior (TSG) algorithm for the selection of the optimal arm, we propose our Thompson Sampling with Clustered arms under Gaussian prior (TSCG) specific to the 2-level hierarchical structure. We prove that by utilizing the 2-level structure, we can achieve a lower regret bound than we do with ordinary TSG. In addition, when the reward is Unimodal, we can reach an even lower bound on the regret by our Unimodal Thompson Sampling algorithm with Clustered Arms under Gaussian prior (UTSCG). Each of our proposed algorithms are accompanied by theoretical evaluation of the upper regret bound, and our numerical experiments confirm the advantage of our proposed algorithms.

Motivation & Objective

  • Identify and formalize a class of optimization problems with clustered Gaussian feedback and a unique optimal arm.
  • Develop algorithms that exploit cluster structure and unimodality to reduce regret.
  • Provide theoretical regret bounds for the proposed algorithms under Gaussian rewards.
  • Demonstrate empirical improvements over baseline methods in simulated mmWave and portfolio tasks.

Proposed method

  • Model arms as Gaussian, partitioned into K clusters with a unique optimal arm.
  • Extend Thompson Sampling with Gaussian priors (TSG) to a two-level structure (TSCG) that selects clusters before arms.
  • further enhance with UTSCG to exploit unimodality inside each cluster by focusing on leaders and neighboring arms.
  • Prove problem-dependent regret bounds for TSG (Theorem 1), TSCG (Theorem 2), and UTSCG (Theorem 3).
  • Assume Strong Dominance and unimodality within clusters to derive tighter bounds.
  • Validate via simulations in mmWave beam/frequency selection and portfolio-style arm ecosystems.

Experimental results

Research questions

  • RQ1How can Gaussian-arm bandits with clustered structure be efficiently learned when the optimal arm is unique?
  • RQ2Can leveraging cluster structure and unimodality reduce regret beyond standard Gaussian-prior Thompson Sampling?
  • RQ3What are the regret bounds for TSCG and UTSCG under Gaussian rewards and unimodal clusters?
  • RQ4Do empirical results in mmWave and portfolio-like settings align with theoretical improvements?

Key findings

  • TSCG achieves lower regret than vanilla TSG by exploiting cluster structure (Theorem 2).
  • UTSCG further reduces regret by leveraging unimodality inside the optimal cluster (Theorem 3).
  • Both algorithms outperform baselines (TSG, UCB, TLP) in cumulative regret and rate of selecting the true optimal arm in simulations.
  • Experiments in mmWave and portfolio scenarios confirm cluster-informed approaches provide earlier convergence to the optimal arm.
  • Theoretical results indicate regret bounds depend on number of clusters, size of the optimal cluster, and clustering quality, while remaining independent of total arm count.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.