Skip to main content
QUICK REVIEW

[논문 리뷰] First-Order Bayesian Regret Analysis of Thompson Sampling

Sébastien Bubeck, Mark Sellke|arXiv (Cornell University)|2019. 02. 02.
Advanced Bandit Algorithms Research인용 수 7
한 줄 요약

이 논문은 조합적 밴딧 문제에 대한 톰슨 샘플링의 정교한 정보이론적 분석을 제안하며, 척도 민감한 정보 비율과 좌표 엔트로피를 도입하여 반복적 반복 설정에서 $\widetilde{O}(\sqrt{dL^*})$의 최초의 베이지안 적 regret 한계를 달성한다. 이는 기존의 최고 수준의 빈도주의 한계와 일치한다. 또한, $L^* \leq \overline{L}^*$일 때 $T$-독립적 regret을 달성하는 Thresholded Thompson Sampling을 제안하며, 이는 표준 톰슨 샘플링이 갖지 못하는 성질이다.

ABSTRACT

We address online combinatorial optimization when the player has a prior over the adversary's sequence of losses. In this framework, Russo and Van Roy proposed an information-theoretic analysis of Thompson Sampling based on the information ratio, resulting in optimal worst-case regret bounds. In this paper we introduce three novel ideas to this line of work. First we propose a new quantity, the scale-sensitive information ratio, which allows us to obtain more refined first-order regret bounds (i.e., bounds of the form $\sqrt{L^*}$ where $L^*$ is the loss of the best combinatorial action). Second we replace the entropy over combinatorial actions by a coordinate entropy, which allows us to obtain the first optimal worst-case bound for Thompson Sampling in the combinatorial setting. Finally, we introduce a novel link between Bayesian agents and frequentist confidence intervals. Combining these ideas we show that the classical multi-armed bandit first-order regret bound $ ilde{O}(\sqrt{d L^*})$ still holds true in the more challenging and more general semi-bandit scenario. This latter result improves the previous state of the art bound $ ilde{O}(\sqrt{(d+m^3)L^*})$ by Lykouris, Sridharan and Tardos. Moreover we sharpen these results with two technical ingredients. The first leverages a recent insight of Zimmert and Lattimore to replace Shannon entropy with more refined potential functions in the analysis. The second is a \emph{Thresholded} Thompson sampling algorithm, which slightly modifies the original algorithm by never playing low-probability actions. This thresholding results in fully $T$-independent regret bounds when $L^*$ is almost surely upper-bounded, which we show does not hold for ordinary Thompson sampling.

연구 동기 및 목표

  • 조합적 반복적 밴딧 문제에서 베이지안과 빈도주의 regret 한계 사이의 격차를 해소하기 위해.
  • 최적의 손실 $L^*$에 따라 변화하는 regret 척도를 반영하는 첫 번째 순서의 regret 스케일링을 포괄하는 정교한 정보이론적 분석을 개발하기 위해.
  • 제한된 최적의 손실 $L^* \leq \overline{L}^*$ 조건 하에서 $T$-독립적 regret 한계를 확립하기 위해. 이는 표준 톰슨 샘플링이 달성하지 못하는 성질이다.
  • 베이지안 에이전트와 빈도주의 신뢰구간을 연결하여, 베이지안과 빈도주의 관점을 통합하기 위해.

제안 방법

  • 손실에 따라 변하는 regret 스케일링을 반영하는 정보 비율의 보완으로 척도 민감한 정보 비율을 도입하기 위해.
  • 조합적 행동에 대한 전역 샤논 엔트로피를 좌표 엔트로피로 대체하여 조합적 설정에서의 분석을 향상시키기 위해.
  • 샤논 엔트로피를 초월하는 정교한 정보이론적 분석을 위해 티슬레스 엔트로피와 로그-바리어 퍼텐셜 함수를 활용하기 위해.
  • 낮은 확률의 행동을 피하는 방식으로 안정성을 확보하며, 제한된 $L^*$ 조건 하에서 $T$-독립적 regret을 달성하기 위해 Thresholded Thompson Sampling을 제안하기 위해.
  • 좋은 암호에 대한 사후 확률에 대한 균일한 하한을 보여주기 위해 새로운 베이지안 가설 검정 기법을 사용하기 위해.
  • 문제를 베이지안 설정으로 축소하기 위해 Sion 최소화 정리를 적용하여, regret 분석에서 사전 분포의 사용을 가능하게 하기 위해.

실험 결과

연구 질문

  • RQ1톰슨 샘플링은 반복적 반복 밴딧 설정에서 $O(\sqrt{dL^*})$ 형태의 첫 번째 순서 regret 한계를 달성할 수 있는가?
  • RQ2표준 톰슨 샘플링 알고리즘이 $L^* \leq \overline{L}^*$일 때 거의 확실히 $T$-독립적 regret을 달성하는가?
  • RQ3톰슨 샘플링의 베이지안 분석을 의미 있는 방식으로 빈도주의 신뢰구간과 연결할 수 있는가?
  • RQ4정보 비율 프레임워크를 척도 민감하거나 좌표 기반 측정으로 정교화하여 regret 한계를 향상시킬 수 있는가?
  • RQ5작은 손실 영역에서 $\Omega(d\overline{L}^*)$의 regret을 피하는 톰슨 샘플링의 변종이 존재하는가?

주요 결과

  • 논문은 반복적 반복 밴딧 설정에서 톰슨 샘플링이 $\widetilde{O}(\sqrt{dL^*})$의 regret을 달성함을 입증하며, 이는 기존의 최고 수준의 빈도주의 한계와 일치한다.
  • 척도 민감한 정보 비율은 손실에 따라 변화하는 스케일링을 반영함으로써 더 엄밀한 첫 번째 순서 regret 분석을 가능하게 한다.
  • 좌표 엔트로피는 전역 샤논 엔트로피를 대체하여, 조합적 설정에서 톰슨 샘플링에 대한 최초의 최악의 경우 regret 한계를 달성한다.
  • Thresholded Thompson Sampling은 $L^* \leq \overline{L}^*$일 때 $T$-독립적 regret을 달성하며, 이는 표준 톰슨 샘플링이 만족하지 못하는 성질이다.
  • 논문은 표준 톰슨 샘플링이 $L^*=0$인 문맥 기반 밴딧에서조차도 $d=O(\sqrt{T})$일지라도 고려할 만한 확률로 $\Omega(\sqrt{T})$의 regret을 겪는다는 것을 증명하며, 이는 작은 손실 영역에서 실패함을 보여준다.
  • 베이지안 에이전트와 빈도주의 신뢰구간 사이에 새로운 연결 고리를 설정하여 통합된 분석 프레임워크를 가능하게 하였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.