Skip to main content
QUICK REVIEW

[论文解读] Estimating Predictability: Redundancy and Surrogate Data Method

Palu\v{s}, M., Ladislav Pecen|arXiv (Cornell University)|Jul 31, 1995
Fault Detection and Control Systems被引用 9
一句话总结

本文提出一种非参数方法,利用信息论冗余性和置换数据技术来估计时间序列的理论可预测性。通过将滞后输入与预测输出之间的冗余度归一化为最大可能冗余度,再与随机打乱(白噪声)和同谱(有色噪声)置换数据进行比较,以分类可预测性为线性、非线性或不可预测,已在合成数据和外汇数据上成功验证。

ABSTRACT

A method for estimating theoretical predictability of time series is presented, based on information-theoretic functionals---redundancies and surrogate data technique. The redundancy, designed for a chosen model and a prediction horizon, evaluates amount of information between a model input (e.g., lagged versions of the series) and a model output (i.e., a series lagged by the prediction horizon from the model input) in number of bits. This value, however, is influenced by a method and precision of redundancy estimation and therefore it is a) normalized by maximum possible redundancy (given by the precision used), and b) compared to the redundancies obtained from two types of the surrogate data in order to obtain reliable classification of a series as either unpredictable or predictable. The type of predictability (linear or nonlinear) and its level can be further evaluated. The method is demonstrated using a numerically generated time series as well as high-frequency foreign exchange data and the theoretical predictability is compared to performance of a nonlinear predictor.

研究动机与目标

  • 开发一种非参数方法,用于估计时间序列的理论可预测性。
  • 在复杂系统中区分线性和非线性可预测性。
  • 提供对时间序列是否不可预测、线性可预测或非线性可预测的可靠分类。
  • 提供一种计算高效的替代方案,以替代训练和测试预测器来评估可预测性。
  • 实现跨数据集的可预测性比较,并与实际预测器性能的相关性。

提出的方法

  • 使用信息论函数计算模型输入(滞后时间序列)与输出(滞后预测时延的时间序列)之间的冗余度,单位为比特。
  • 通过熵推导出的最大可能冗余度对冗余度进行归一化,将可预测性表示为理论最大值的百分比。
  • 生成两类置换数据:随机打乱置换数据(代表白噪声)和同谱置换数据(代表线性相关过程)。
  • 计算原始数据及两类置换数据的冗余度,以检验差异的统计显著性。
  • 根据其冗余度是否显著高于随机打乱或同谱置换数据,对序列进行分类。
  • 采用滑动窗口方法,窗口大小 Nw = 256,步长 Ns = 4,以估计局部可预测性。

实验结果

研究问题

  • RQ1基于冗余度的度量能否在不训练预测器的情况下可靠估计时间序列的理论可预测性?
  • RQ2该方法能否在时间序列中区分线性和非线性可预测性?
  • RQ3置换数据类型(随机打乱 vs. 同谱)在多大程度上提高了可预测性分类的可靠性?
  • RQ4理论可预测性与实际非线性预测器性能的相关性如何?
  • RQ5该方法能否检测非平稳序列中可预测性结构随时间的变化?

主要发现

  • 对于合成的 USD/JPY 时间序列,该方法正确识别出后半段非线性可预测性上升,非线性可预测性从约 2 比特增至约 4 比特,而线性可预测性保持稳定。
  • 在 GBP/USD 时间序列中,预测误差与线性可预测性指标强相关,但与非线性或总可预测性无关,表明该时段以线性结构为主导。
  • 该方法成功将合成数据分类为:前半段为线性可预测,后半段为非线性可预测,与数据生成过程一致。
  • 理论可预测性度量与高频外汇数据中样条非线性预测器的性能高度对应,尽管结果尚属初步。
  • 置换比较方法有效过滤了虚假信号,随机打乱与同谱置换数据之间冗余度的差异为分类提供了统计置信度。
  • 该方法的计算成本显著低于训练和测试任何预测器,使其在初步分析中具有高效性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。