Skip to main content
QUICK REVIEW

[논문 리뷰] Fast Online Learning with Gaussian Prior-Driven Hierarchical Unimodal Thompson Sampling

Tianchi Zhao, He Liu|arXiv (Cornell University)|2026. 02. 17.
Advanced Bandit Algorithms Research인용 수 0
한 줄 요약

본 논문은 Gaussian-arm 밴디트에 대해 군집화된 구조와 단모드 구조를 갖는 두 가지 Thompson Sampling 변형(TSCG 및 UTSCG)을 제시하고, 더 촘촘한 후회(bound)를 증명하고 mmWave 및 포트폴리오 시나리오에서 성능을 검증한다.

ABSTRACT

We study a type of Multi-Armed Bandit (MAB) problems in which arms with a Gaussian reward feedback are clustered. Such an arm setting finds applications in many real-world problems, for example, mmWave communications and portfolio management with risky assets, as a result of the universality of the Gaussian distribution. Based on the Thompson Sampling algorithm with Gaussian prior (TSG) algorithm for the selection of the optimal arm, we propose our Thompson Sampling with Clustered arms under Gaussian prior (TSCG) specific to the 2-level hierarchical structure. We prove that by utilizing the 2-level structure, we can achieve a lower regret bound than we do with ordinary TSG. In addition, when the reward is Unimodal, we can reach an even lower bound on the regret by our Unimodal Thompson Sampling algorithm with Clustered Arms under Gaussian prior (UTSCG). Each of our proposed algorithms are accompanied by theoretical evaluation of the upper regret bound, and our numerical experiments confirm the advantage of our proposed algorithms.

연구 동기 및 목표

  • Identify and formalize a class of optimization problems with clustered Gaussian feedback and a unique optimal arm.
  • Develop algorithms that exploit cluster structure and unimodality to reduce regret.
  • Provide theoretical regret bounds for the proposed algorithms under Gaussian rewards.
  • Demonstrate empirical improvements over baseline methods in simulated mmWave and portfolio tasks.

제안 방법

  • 모든 팔을 Gaussian으로 모델링하고, 고유한 최적 팔을 가진 K개의 군집으로 분할한다.
  • Extend Thompson Sampling with Gaussian priors (TSG) to a two-level structure (TSCG) that selects clusters before arms.
  • further enhance with UTSCG to exploit unimodality inside each cluster by focusing on leaders and neighboring arms.
  • Prove problem-dependent regret bounds for TSG (Theorem 1), TSCG (Theorem 2), and UTSCG (Theorem 3).
  • Assume Strong Dominance and unimodality within clusters to derive tighter bounds.
  • Validate via simulations in mmWave beam/frequency selection and portfolio-style arm ecosystems.

실험 결과

연구 질문

  • RQ1How can Gaussian-arm bandits with clustered structure be efficiently learned when the optimal arm is unique?
  • RQ2Can leveraging cluster structure and unimodality reduce regret beyond standard Gaussian-prior Thompson Sampling?
  • RQ3What are the regret bounds for TSCG and UTSCG under Gaussian rewards and unimodal clusters?
  • RQ4Do empirical results in mmWave and portfolio-like settings align with theoretical improvements?

주요 결과

  • TSCG achieves lower regret than vanilla TSG by exploiting cluster structure (Theorem 2).
  • UTSCG further reduces regret by leveraging unimodality inside the optimal cluster (Theorem 3).
  • Both algorithms outperform baselines (TSG, UCB, TLP) in cumulative regret and rate of selecting the true optimal arm in simulations.
  • Experiments in mmWave and portfolio scenarios confirm cluster-informed approaches provide earlier convergence to the optimal arm.
  • Theoretical results indicate regret bounds depend on number of clusters, size of the optimal cluster, and clustering quality, while remaining independent of total arm count.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.