[论文解读] Learning Dynamical Systems via Koopman Operator Regression in Reproducing Kernel Hilbert Spaces
本文提出了一种在再生核希尔伯特空间(RKHS)中进行Koopman算子回归的统计学习框架,通过核方法实现对动力系统的数据驱动建模。该研究引入了一种低秩估计器(RRR),其非渐近学习界在混合条件下成立,并在空气质量与分子动力学数据上的实验中,展示了其在预测与模态分解方面优于基线方法的性能。
We study a class of dynamical systems modelled as Markov chains that admit an invariant distribution via the corresponding transfer, or Koopman, operator. While data-driven algorithms to reconstruct such operators are well known, their relationship with statistical learning is largely unexplored. We formalize a framework to learn the Koopman operator from finite data trajectories of the dynamical system. We consider the restriction of this operator to a reproducing kernel Hilbert space and introduce a notion of risk, from which different estimators naturally arise. We link the risk with the estimation of the spectral decomposition of the Koopman operator. These observations motivate a reduced-rank operator regression (RRR) estimator. We derive learning bounds for the proposed estimator, holding both in i.i.d. and non i.i.d. settings, the latter in terms of mixing coefficients. Our results suggest RRR might be beneficial over other widely used estimators as confirmed in numerical experiments both for forecasting and mode decomposition.
研究动机与目标
- 将动力系统有限轨迹的Koopman算子回归形式化为统计学习框架。
- 将Koopman算子的估计与RKHS中的风险最小化及谱分解联系起来。
- 开发并分析一种新型低秩估计器(RRR),以提升泛化能力与可解释性。
- 基于混合系数,在独立同分布(i.i.d.)与非独立同分布(non-i.i.d.)设定下,推导RRR估计器的非渐近学习界。
- 通过在预测与模态分解任务上的实证验证,展示该框架的有效性。
提出的方法
- 将Koopman算子限制在再生核希尔伯特空间(RKHS)中,使希尔伯特-施密特算子可作为假设空间使用。
- 将Koopman算子回归表述为风险最小化问题,其风险与谱分解误差相关联。
- 通过约束Koopman算子估计的秩,引入低秩算子回归(RRR)估计器。
- 利用混合过程的测度集中性理论推导学习界,将经典的i.i.d.结果推广至非i.i.d.时间序列。
- 采用核方法,包括线性核、高斯核以及基于深度神经网络的核(如CNN嵌入),将状态观测映射至RKHS。
- 在实验中使用时间序列交叉验证(TimeSeriesSplit)与网格搜索进行正则化参数选择。
实验结果
研究问题
- RQ1如何在RKHS中将Koopman算子回归形式化为统计学习框架?
- RQ2Koopman算子估计器的风险与其谱分解精度之间有何关系?
- RQ3低秩估计器(RRR)在从有限、非i.i.d.轨迹中学习动力系统时,如何提升泛化能力与性能?
- RQ4在弱依赖性(混合)假设下,能否为Koopman算子回归建立非渐近学习界?
- RQ5在多种真实世界数据集上,基于核的Koopman估计器在预测与模态分解中的表现如何?
主要发现
- 在北京市空气质量数据集上,RRR估计器的训练误差与测试误差均低于PCR及其他基线方法,正则化参数γ=10⁻⁴通过网格搜索选定。
- 模态分解揭示了PM2.5浓度峰值与后续风速峰值之间存在稳定的2小时延迟,支持关于污染扩散的物理解释。
- 基于CNN的核在长时预测中保持了强劲的预测性能,而线性核与高斯核则迅速退化,表明深度特征学习的优势。
- 理论分析建立了非i.i.d.设定下RRR估计器的非渐近学习界,其依赖于关于混合过程的新型引理。
- 该框架可将动态模态分解(DMD)及相关方法作为特例恢复,从统计学习视角统一了这些方法。
- 数值实验确认,RRR估计器在预测精度与模态分解质量方面均优于标准估计器。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。