Skip to main content
QUICK REVIEW

[论文解读] Regularized Estimation in High Dimensional Time Series under Mixing Conditions.

Kam Chung Wong, Ambuj Tewari|arXiv (Cornell University)|Feb 12, 2016
Statistical Methods and Inference参考文献 14被引用 4
一句话总结

该论文在弱依赖性时间序列的混合条件假设下,为高维时间序列中的Lasso建立了非渐近误差界,无需依赖高斯性或有限阶VAR结构等参数假设。论文提出了一类适用于β-混合子高斯向量的新型Hanson-Wright不等式,使得Lasso在非高斯和非线性时间序列模型中仍能保持一致性。

ABSTRACT

The Lasso is one of the most popular methods in high dimensional statistical learning. Most existing theoretical results for the Lasso, however, require the samples to be iid. Recent work has provided guarantees for the Lasso assuming that the time series is generated by a sparse Vector Auto-Regressive (VAR) model with Gaussian innovations. Proofs of these results rely critically on the fact that the true data generating mechanism (DGM) is a finite-order Gaussian VAR. This assumption is quite brittle: linear transformations, including selecting a subset of variables, can lead to the violation of this assumption. In order to break free from such assumptions, we derive nonasymptotic inequalities for estimation error and prediction error of the Lasso estimate of the best linear predictor without assuming any special parametric form of the DGM. Instead, we rely only on (strict) stationarity and mixing conditions to establish consistency of the Lasso in the following two scenarios: (a) alpha-mixing Gaussian processes, and (b) beta-mixing sub-Gaussian random vectors. Our work provides an alternative proof of the consistency of the Lasso for sparse Gaussian VAR models. But the applicability of our results extends to non-Gaussian and non-linear times series models as the examples we provide demonstrate. In order to prove our results, we derive a novel Hanson-Wright type concentration inequality for beta-mixing sub-Gaussian random vectors that may be of independent interest.

研究动机与目标

  • 放松现有高维时间序列Lasso理论中对严格参数假设(如有限阶高斯VAR模型)的依赖。
  • 在仅假设平稳性和混合条件的前提下,建立Lasso的非渐近估计与预测误差界。
  • 通过依赖弱依赖性而非特定参数化的DGM,将Lasso的一致性推广至非高斯和非线性时间序列模型。
  • 为β-混合子高斯随机向量开发一种新的浓度不等式,其适用范围超越现有结果。
  • 提供一个统一的理论框架,使高维时间序列中Lasso的理论分析对线性变换和模型误设具有鲁棒性。

提出的方法

  • 仅基于(严格)平稳性和α-或β-混合条件,推导Lasso的非渐近风险界。
  • 提出一种专为β-混合子高斯随机向量设计的新型Hanson-Wright型浓度不等式。
  • 将该浓度不等式应用于控制高维回归中的经验过程与设计矩阵偏差。
  • 在弱依赖性下,建立Lasso对最优线性预测器的一致性,且不假设真实DGM具有特定参数形式。
  • 证明当真实数据生成机制为非高斯或非线性时,Lasso仍保持一致性。
  • 通过链式论证和矩量界控制混合条件下经验过程的上确界。

实验结果

研究问题

  • RQ1在不假设有限阶高斯VAR模型的前提下,Lasso能否在高维时间序列中实现一致估计?
  • RQ2当底层时间序列为严格平稳且混合时,Lasso的非渐近误差界是什么?
  • RQ3如何将浓度不等式扩展至在β-混合条件下依赖的子高斯随机向量?
  • RQ4在弱依赖性下,Lasso对非高斯和非线性时间序列模型是否仍保持一致性?
  • RQ5能否仅通过混合条件将Lasso的理论保证扩展至参数模型之外?

主要发现

  • 在仅假设α-混合高斯过程且无需参数化DGM的前提下,Lasso实现了非渐近估计与预测误差界。
  • 在β-混合子高斯随机向量下,Lasso对最优线性预测器具有一致性,即使真实DGM为非高斯或非线性。
  • 为β-混合子高斯向量推导出一种新型Hanson-Wright型浓度不等式,该不等式在其他高维依赖数据场景中亦具应用潜力。
  • 该理论框架可作为稀疏高斯VAR模型下已有结果的特例,且在更弱假设下提供了替代证明。
  • 结果对数据的线性变换具有鲁棒性,而此类变换可能破坏有限阶VAR结构等参数假设。
  • 该方法适用于比以往Lasso理论更广泛的时间序列模型类别,包括在弱依赖性下的非线性和非高斯过程。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。