[论文解读] Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance Sampling
本文对通过条件重要性采样进行离策略评估的理论分析表明,在有限时域马尔可夫决策过程(MDP)中,每决策重要性采样(PDIS)和稳态重要性采样(SIS)并不总是能降低与原始重要性采样(vanilla IS)相比的方差。令人惊讶的是,它证明了在某些条件下,这些估计器的方差可能比粗略IS高出指数级,但渐近意义上,它们可实现多项式方差增长,从而解释了其在长时域设置中的经验成功。
Off-policy policy estimators that use importance sampling (IS) can suffer from high variance in long-horizon domains, and there has been particular excitement over new IS methods that leverage the structure of Markov decision processes. We analyze the variance of the most popular approaches through the viewpoint of conditional Monte Carlo. Surprisingly, we find that in finite horizon MDPs there is no strict variance reduction of per-decision importance sampling or stationary importance sampling, comparing with vanilla importance sampling. We then provide sufficient conditions under which the per-decision or stationary estimators will provably reduce the variance over importance sampling with finite horizons. For the asymptotic (in terms of horizon $T$) case, we develop upper and lower bounds on the variance of those estimators which yields sufficient conditions under which there exists an exponential v.s. polynomial gap between the variance of importance sampling and that of the per-decision or stationary estimators. These results help advance our understanding of if and when new types of IS estimators will improve the accuracy of off-policy estimation.
研究动机与目标
- 理解为何在理论上预期方差降低的情况下,稳态重要性采样(SIS)和每决策重要性采样(PDIS)在离策略评估中有时仍无法降低方差。
- 通过扩展条件蒙特卡洛估计器的视角,形式化IS、PDIS和SIS之间的关系。
- 推导出在有限时域MDP中,PDIS和SIS相比原始IS可保证降低方差的充分条件。
- 建立IS、PDIS和SIS方差的渐近上下界,表明其对时域T的依赖关系为多项式或指数级。
- 通过识别方差降低的条件,将理论发现与SIS在长时域领域中的经验成功相协调。
提出的方法
- 通过基于不同统计量(如状态-动作序列、奖励)的条件化,将IS、PDIS和SIS形式化为条件蒙特卡洛估计器的实例。
- 应用全方差定律分析条件估计器的方差,表明由于轨迹和中各项之间的协方差,方差降低并非必然。
- 在定理1和定理2中,基于MDP结构和条件化统计量的选择,推导出PDIS和SIS方差降低的充分条件。
- 利用鞅的集中不等式,推导出在一般状态空间下所有三种估计器的渐近方差界。
- 比较渐近方差增长速率:表明在某些条件下,SIS和PDIS对时域T具有多项式(如O(T²))依赖,而原始IS则具有指数依赖。
- 证明在批量设置中,基于奖励的估计器等价于原始IS,凸显在线学习与批量学习之间的差异。
实验结果
研究问题
- RQ1在有限时域MDP中,每决策重要性采样(PDIS)在何种条件下相比原始重要性采样(IS)能降低方差?
- RQ2稳态重要性采样(SIS)是否可被证明在方差上优于PDIS和IS?如果是,其充分条件是什么?
- RQ3为何经验结果表明SIS在长时域领域中优于IS,尽管其理论方差存在更高风险?
- RQ4随着时域T的增长,IS、PDIS和SIS的渐近方差行为如何?它们在增长速率上如何比较?
- RQ5对不同统计量(如奖励、状态-动作对)进行条件化,如何影响离策略评估中重要性采样估计器的方差?
主要发现
- 在有限时域MDP中,每决策重要性采样和稳态重要性采样并不总是能降低相比原始IS的方差;存在反例表明IS的方差更低。
- 由于轨迹和中各项之间的协方差,条件重要性采样估计器的方差并不保证低于粗略IS。
- 定理1和定理2中推导出PDIS和SIS方差降低的充分条件,基于MDP结构和条件化统计量的选择。
- 渐近意义上,在某些条件下,SIS的方差增长为O(T²),而原始IS的方差可能随T呈指数增长,从而解释了SIS在长时域设置中的经验成功。
- 推论4表明,在弱假设下,SIS的方差可被证明低于PDIS,为SIS在某些场景中的优越性提供了理论支持。
- 在批量设置中,基于奖励的估计器被证明等价于原始IS,表明此类条件化在批量离策略评估中并不必然降低方差。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。