[论文解读] On the Relation between Realizable and Nonrealizable Cases of the Sequence Prediction Problem
本文建立了在不同损失函数下序列预测的可实现与不可实现情形之间的正式关系。结果表明,当以总变差距离衡量预测质量时,两种情形等价;但在期望平均KL散度下则不等价。主要贡献在于证明了两种问题的解均可构造为给定测度类中可数子集上的贝叶斯混合,且对特定类(如有限记忆过程和平稳过程)完全刻画了解存在的条件。
A sequence $x_1,\dots,x_n,\dots$ of discrete-valued observations is generated according to some unknown probabilistic law (measure) $μ$. After observing each outcome, one is required to give conditional probabilities of the next observation. The realizable case is when the measure $μ$ belongs to an arbitrary but known class $\mathcal C$ of process measures. The non-realizable case is when $μ$ is completely arbitrary, but the prediction performance is measured with respect to a given set $\mathcal C$ of process measures. We are interested in the relations between these problems and between their solutions, as well as in characterizing the cases when a solution exists and finding these solutions. We show that if the quality of prediction is measured using the total variation distance, then these problems coincide, while if it is measured using the expected average KL divergence, then they are different. For some of the formalizations we also show that when a solution exists, it can be obtained as a Bayes mixture over a countable subset of $\mathcal C$. We also obtain several characterization of those sets $\mathcal C$ for which solutions to the considered problems exist. As an illustration to the general results obtained, we show that a solution to the non-realizable case of the sequence prediction problem exists for the set of all finite-memory processes, but does not exist for the set of all stationary processes. It should be emphasized that the framework is completely general: the processes measures considered are not required to be i.i.d., mixing, stationary, or to belong to any parametric family.
研究动机与目标
- 澄清可实现与不可实现序列预测情形之间的正式关系。
- 确定在真实测度 μ 不一定属于类 C 的情况下,预测器存在的条件。
- 刻画在可实现与不可实现设置下解存在的集合 C。
- 证明当解存在时,其可构造为 C 的可数子集上的贝叶斯混合。
- 通过具体例子(如有限记忆过程和平稳过程)说明理论。
提出的方法
- 本文使用总变差距离和期望平均KL散度作为损失函数,比较可实现与不可实现预测问题。
- 通过使用正则化项 γ 确保适当归一化与收敛性,将预测器 ν 构造为 C 中测度的加权混合。
- 该方法涉及基于似然比定义阈值集 Tⁿⱼₙᵤ,并利用熵界控制 μ 与 ν 之间的分歧。
- 关键技术步骤是将分歧 dₙ(μ,ν) 分解为两部分:I(在高似然集上)与 II(在其余部分),并证明两者均为 o(n)。
- 正则化项 γ 被替换为 C 中可数个测度的凸组合,以确保预测器保持可计算且定义良好。
- 证明依赖于构造权重序列 wₙ,并利用 −log εₙᵤ = o(n) 的事实,以确保分歧项的收敛性。
实验结果
研究问题
- RQ1在何种条件下,可实现与不可实现序列预测问题等价?
- RQ2损失函数的选择(总变差距离 vs. 期望平均KL散度)如何影响两个问题的等价性?
- RQ3对于哪些过程测度类 C,不可实现情形下存在预测器?
- RQ4是否可将序列预测问题的解构造为 C 的可数子集上的贝叶斯混合?
- RQ5在可实现与不可实现设置下,预测器存在的必要与充分条件是什么?
主要发现
- 当损失以总变差距离衡量时,可实现与不可实现的序列预测情形等价。
- 当损失以期望平均KL散度衡量时,两种情形本质上不同。
- 只要解存在,两种问题的解均可构造为 C 的可数子集上的贝叶斯混合。
- 当 C 为所有有限记忆过程的集合时,不可实现情形下存在解。
- 当 C 为所有平稳过程的集合时,不可实现情形下不存在解。
- 在所构造的预测器 ν 下,分歧 dₙ(μ,ν) 以平均意义收敛至零(即 dₙ(μ,ν)/n → 0),证实了渐近一致性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。