Skip to main content
QUICK REVIEW

[논문 리뷰] Non-Stationary Bandits with Habituation and Recovery Dynamics

Yonatan Mintz, Anil Aswani|arXiv (Cornell University)|2017. 07. 26.
Advanced Bandit Algorithms Research참고 문헌 80인용 수 12
한 줄 요약

이 논문은 순차적 결정 문제에서 습관화 및 회복 동역학을 모델링하는 새로운 비정상적 다수의 손잡이 밴딧 프레임워크인 ROGUE 밴딧을 소개한다. 최대우도추정과 유한표본 농도 경계를 활용하여 저자들은 ROGUE-UCB 및 ε-ROGUE 알고리즘을 개발하였으며, 이는 로그적 감소를 달성하여 개인 맞춤형 헬스케어 간섭에서 최신 기술보다 뛰어나며 시뮬레이션에서 하루 평균 보행 수를 약 1,000보 증가시킨다.

ABSTRACT

Many settings involve sequential decision-making where a set of actions can be chosen at each time step, each action provides a stochastic reward, and the distribution for the reward of each action is initially unknown. However, frequent selection of a specific action may reduce its expected reward, while abstaining from choosing an action may cause its expected reward to increase. Such non-stationary phenomena are observed in many real world settings such as personalized healthcare-adherence improving interventions and targeted online advertising. Though finding an optimal policy for general models with non-stationarity is PSPACE-complete, we propose and analyze a new class of models called ROGUE (Reducing or Gaining Unknown Efficacy) bandits, which we show in this paper can capture these phenomena and are amenable to the design of effective policies. We first present a consistent maximum likelihood estimator for the parameters of these models. Next, we construct finite sample concentration bounds that lead to an upper confidence bound policy called the ROGUE Upper Confidence Bound (ROGUE-UCB) algorithm. We prove that under proper conditions the ROGUE-UCB algorithm achieves logarithmic in time regret, unlike existing algorithms which result in linear regret. We conclude with a numerical experiment using real data from a personalized healthcare-adherence improving intervention to increase physical activity. In this intervention, the goal is to optimize the selection of messages (e.g., confidence increasing vs. knowledge increasing) to send to each individual each day to increase adherence and physical activity. Our results show that ROGUE-UCB performs better in terms of regret and average reward as compared to state of the art algorithms, and the use of ROGUE-UCB increases daily step counts by roughly 1,000 steps a day (about a half-mile more of walking) as compared to other algorithms.

연구 동기 및 목표

  • 반복적인 행동 사용이 습관화를 유도하고, 중단 시 회복이 일어나는 행동 설정에서 정상적 밴딧 모델의 한계를 해결하기 위해.
  • 행동 빈도가 장기적 유효성에 영향을 주는 개인 맞춤형 헬스케어 간섭에서 비정상적 보상 동역학을 모델링하기 위해.
  • 보상 동역학에 대한 사전 지식이 필요 없이 습관화 및 회복을 포괄하는 탄력적이고 분석 가능한 밴딧 프레임워크를 개발하기 위해.
  • 기존 방법이 실패하는 비정상 환경에서 증명 가능한 감소 보장을 갖춘 알고리즘을 설계하기 위해.
  • 실제 신체 활동 간섭 데이터를 활용하여 제안된 알고리즘의 경험적 우수성을 입증하기 위해.

제안 방법

  • 행동 보상에서 습관화 및 회복을 모델링하기 위해 새로운 비정상적 밴딧 클래스인 ROGUE(알려지지 않은 유효성 감소 또는 증가) 밴딧을 제안한다.
  • 관측된 보상과 행동 역사를 기반으로 모델 파라미터를 학습하기 위한 일致한 최대우도추정(MLE) 방법을 개발한다.
  • 알고리즘 성능의 엄밀한 이론적 분석을 가능하게 하기 위해 MLE 추정치에 대한 유한표본 농도 경계를 유도한다.
  • 추정된 불확실성에 기반해 탐색과 이용을 균형 잡는 상한 신뢰도 기반 알고리즘인 ROGUE-UCB를 설계한다.
  • 감소하는 탐색 비율을 갖는 ε-그리디 알고리즘인 ε-ROGUE를 설계한다.
  • 오프라인 파라미터 추정을 위한 혼합정수계획법 재구성 가능성을 보장하기 위해 보상 가능성에 라플라스 노이즈 모델을 사용한다.

실험 결과

연구 질문

  • RQ1비정상적 밴딧 모델은 행동 간섭에서 습관화 및 회복 동역학을 효과적으로 포착할 수 있는가?
  • RQ2비정상성에도 불구하고 로그적 감소를 달성할 수 있는 증명 가능한 효율적 알고리즘이 ROGUE 밴딧에 대해 설계될 수 있는가?
  • RQ3MLE 기반의 파라미터 추정과 유한표본 농도 경계는 이 설정에서 신뢰할 수 있는 정책 학습을 어떻게 가능하게 하는가?
  • RQ4ROGUE-UCB와 ε-ROGUE는 감소와 평균 보상 측면에서 기존 최신 기술 밴딧 알고리즘을 초월하는가?
  • RQ5제안된 프레임워크는 실생활 개인 맞춤형 헬스케어 간섭에서 결과를 크게 향상시킬 수 있는가?

주요 결과

  • ROGUE-UCB와 ε-ROGUE는 기존 알고리즘이 비정상적 환경에서 선형 감소를 겪는 것과는 달리 기대값에서 로그적 감소를 달성한다.
  • 개인 맞춤형 신체 활동 간섭에서의 실세계 데이터를 활용한 시뮬레이션에서, ROGUE-UCB는 다음으로 우수한 알고리즘 대비 하루 평균 보행 수를 약 1,000보 증가시켰다.
  • 초기 단계에서의 탐색 감소로 인해 ROGUE-UCB는 누적 감소 측면에서 ε-ROGUE를 능가했으며, 초기 보상도 더 높았다.
  • ROGUE-UCB와 ε-ROGUE는 순수 탐색 및 D-UCB에 비해 평균 보상과 누적 감소 측면에서 유의미하게 뛰어났다.
  • MLE 추정치에 대한 유한표본 농도 경계가 성공적으로 유도되었으며, 이는 제안된 알고리즘의 이론적 성능 보장을 정당화하는 데 사용되었다.
  • 보상 모델에 라플라스 노이즈를 사용함으로써 혼합정수계획법 솔버를 통한 효과적인 오프라인 파라미터 추정이 가능해졌으며, 실용적 구현에 기여했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.