Skip to main content
QUICK REVIEW

[논문 리뷰] An Exponential Lower Bound for Linearly-Realizable MDPs with Constant Suboptimality Gap

Yuanhao Wang, Ruosong Wang|arXiv (Cornell University)|2021. 03. 23.
Advanced Bandit Algorithms Research참고 문헌 44인용 수 8
한 줄 요약

이 논문은 일정한 부분최적성 갭을 가진 선형적으로 실현 가능한 MDP에서 강화학습의 지수적 샘플 복잡도 하한을 확립한다. 이는 조건이 유리한 갭 가정이 있더라도 추가적인 구조적 제약 조건이 없으면 온라인 RL이 여전히 비효율적임을 보여주며, 온라인 RL과 생성 모델 설정 사이에 지수적 분리가 존재함을 밝힌다. 이 결과는 동일한 조건 하에서 생성 모델 설정에서는 다항 샘플 복잡도가 달성 가능하다는 점을 시사한다.

ABSTRACT

A fundamental question in the theory of reinforcement learning is: suppose the optimal $Q$-function lies in the linear span of a given $d$ dimensional feature mapping, is sample-efficient reinforcement learning (RL) possible? The recent and remarkable result of Weisz et al. (2020) resolved this question in the negative, providing an exponential (in $d$) sample size lower bound, which holds even if the agent has access to a generative model of the environment. One may hope that this information theoretic barrier for RL can be circumvented by further supposing an even more favorable assumption: there exists a \emph{constant suboptimality gap} between the optimal $Q$-value of the best action and that of the second-best action (for all states). The hope is that having a large suboptimality gap would permit easier identification of optimal actions themselves, thus making the problem tractable; indeed, provided the agent has access to a generative model, sample-efficient RL is in fact possible with the addition of this more favorable assumption. This work focuses on this question in the standard online reinforcement learning setting, where our main result resolves this question in the negative: our hardness result shows that an exponential sample complexity lower bound still holds even if a constant suboptimality gap is assumed in addition to having a linearly realizable optimal $Q$-function. Perhaps surprisingly, this implies an exponential separation between the online RL setting and the generative model setting. Complementing our negative hardness result, we give two positive results showing that provably sample-efficient RL is possible either under an additional low-variance assumption or under a novel hypercontractivity assumption (both implicitly place stronger conditions on the underlying dynamics model).

연구 동기 및 목표

  • 일정한 부분최적성 갭을 가정할 경우 선형적으로 실현 가능한 MDP에서 지수적 샘플 복잡도 하한이 유지되는지 조사하기.
  • Weisz 등(2020)이 온라인 RL 설정에서 규명한 딱지 장벽을 극복하기 위해 부분최적성 갭 가정만으로는 충분한지 판단하기.
  • 낮은 분산 또는 초수렴성과 같은 추가적인 구조적 가정이 샘플 효율성을 회복하는 데 기여하는지 탐색하기.
  • 선형 실현 가능성과 부분최적성 갭 조건 하에서 온라인 RL과 생성 모델 설정 간의 근본적 격차를 명확히 하기.
  • 선형 실현 가능성과 일정한 갭을 초월해 샘플 효율적 RL을 위한 필수 조건을 설정하기.

제안 방법

  • 선형적으로 실현 가능한 최적 Q-함수와 일정한 부분최적성 갭을 가진 MDP의 가족을 구성하여 딱지의 본질을 입증한다.
  • 어려운 밴딧 문제에서의 감소를 통해 어떤 알고리즘도 특성 차원 d에 대해 지수적으로 많은 샘플이 필요하다는 것을 보여준다.
  • 초수렴성 하에서 최소제곱 회귀의 새로운 분석을 통해 이 조건 하에서 양면 결과를 도출한다.
  • Du 등(2019c)과 Weisz 등(2020)의 기법을 적응적으로 활용하며, 유니온 바운드 대신 Lemma 6를 사용해 딱지 MDP 구성에서 추정 오차를 통제한다.
  • 두 가지 양면 결과를 도입한다: 하나는 낮은 분산 가정 하에서, 다른 하나는 역동성에 대한 새로운 초수렴성 조건 하에서.
  • 하한 분석에서 사용된 딱지 MDP 가족이 낮은 분산 및 초수렴성 가정을 모두 위반함을 보여, 이들이 효율성에 필수적임을 시사한다.

실험 결과

연구 질문

  • RQ1온라인 RL 설정에서 일정한 부분최적성 갭을 가정할 경우 선형적으로 실현 가능한 MDP에 대한 지수적 샘플 복잡도 하한이 유지되는가?
  • RQ2생성 모델에 접근할 수 없을 때 부분최적성 갭 가정만으로 온라인 RL에서 샘플 효율적 학습을 보장할 수 있는가?
  • RQ3선형 실현 가능성과 일정한 갭 조건 하에서 온라인 RL에서 다항 샘플 복잡도를 달성하기 위해 추가로 필요한 구조적 가정은 무엇인가?
  • RQ4동일한 가정 하에서 온라인 RL 설정과 생성 모델 설정 간에 지수적 분리가 존재하는가?
  • RQ5낮은 분산 또는 초수렴성 가정은 MDP 복잡성의 천연 특성화로 해석될 수 있는가?

주요 결과

  • 선형 실현 가능성과 일정한 부분최적성 갭이 존재하는 온라인 RL에 대해 $2^{\tilde{\theta}(\text{min}\nolimits\{d,H\})}$ 의 지수적 샘플 복잡도 하한이 성립한다.
  • 생성 모델이 없을 경우에도 하한이 유지되며, 동일한 가정 하에서 온라인 설정이 생성 모델 설정보다 지수적으로 더 어렵다는 것을 보여준다.
  • 낮은 분산 가정(가정 3) 하에서는 샘플 효율적 학습이 가능하지만, 딱지 MDP 가족에서는 이 가정에 필요한 상수 $C$ 가 지수적으로 크다.
  • 초수렴성 가정(가정 4) 하에서는 샘플 효율적 학습이 가능하지만, 딱지 MDP 가족에서는 초수렴성 상수 $C_{\text{hyper}}$ 가 지수적으로 크다.
  • 하한 분석에서 사용된 딱지 MDP 가족은 지수적으로 작은 최소 도달 확률을 가지지만, 최소 도달 확률 $\eta_{\text{min}} = 1$ 으로 수정 가능하며, 이 경우에도 지수적 하한이 유지된다.
  • 결과적으로 낮은 분산 또는 초수렴성 가정이 선형 실현 가능성과 일정한 부분최적성 갭에 의해 유도되지 않음을 보여, 이들이 효율성에 필수적임을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.