Skip to main content
QUICK REVIEW

[论文解读] Are We Winning the Wrong Game? Revisiting Evaluation Practices for Long-Term Time Series Forecasting

Thanapol Phungtua-eng, Yoshitaka Yamamoto|arXiv (Cornell University)|Mar 9, 2026
Forecasting Techniques and Applications被引用 0
一句话总结

该论文认为长期时间序列 forecasting 评估过于依赖聚合的逐点指标(MSE/MAE),提出一个多维诊断评估框架,聚焦统计保真性、结构一致性和决策相关性,将评估重心从排行榜追逐转向情境感知的预测。

ABSTRACT

Long-term time series forecasting (LTSF) is widely recognized as a central challenge in data mining and machine learning. LTSF has increasingly evolved into a benchmark-driven ''GAME,'' where models are ranked, compared, and declared state-of-the-art based primarily on marginal reductions in aggregated pointwise error metrics such as MSE and MAE. Across a small set of canonical datasets and fixed forecasting horizons, progress is communicated through leaderboard-style tables in which lower numerical scores define success. In this GAME, what is measured becomes what is optimized, and incremental error reduction becomes the dominant currency of advancement. We argue that this metric-centric regime is not merely incomplete, but structurally misaligned with the broader objectives of forecasting. In real-world settings, forecasting often prioritizes preserving temporal structure, trend stability, seasonal coherence, robustness to regime shifts, and supporting downstream decision processes. Optimizing aggregate pointwise error does not necessarily imply modeling these structural properties. As a result, leaderboard improvement may increasingly reflect specialization in benchmark configurations rather than a deeper understanding of temporal dynamics. This paper revisits LTSF evaluation as a foundational question in data science: what does it mean to measure forecasting progress? We propose a multi-dimensional evaluation perspective that integrates statistical fidelity, structural coherence, and decision-level relevance. By challenging the current metric monoculture, we aim to redirect attention from winning benchmark tables toward advancing meaningful, context-aware forecasting.

研究动机与目标

  • 挑战当前以聚合逐点误差为焦点的 LTSF 评估实践的充分性。
  • 倡导包含统计保真性、结构一致性与决策级相关性的多维评估框架。
  • 鼓励诊断性报告与窗口级分析,以揭示在结构条件下的性能表现。
  • 倡导以领域特定、情境感知的评估胜过普遍的排行榜支配。

提出的方法

  • 批判性分析现有以基准为驱动的 LTSF 评估机制及其激励。
  • 提出三维评估框架:统计保真性、结构一致性和决策级相关性。
  • 建议诊断性报告与窗口级分析,以评估在不同机制下的性能变动。
  • 主张强调领域特定目标和有条件的表现,而非普遍的领导者。
  • 参考并构建新兴的基准化倡议(如 TSFM-Bench、TIME、TFB)以展示多样化的评估实践。

实验结果

研究问题

  • RQ1当前 LTSF 的评估游戏是否与超越边际 MSE/MAE 降低的真实预测目标一致?
  • RQ2如何将 LTSF 评估重新框定,以平衡统计保真性、时间结构和决策有用性?
  • RQ3哪些实际的评估做法(如窗口级诊断)可以揭示模型在不同结构条件下的行为?
  • RQ4是否存在跨数据集与时间 horizon 的普遍冠军,还是应将评估情境化以契合领域目标?

主要发现

  • 排行榜驱动的进展往往反映基准特定的适应性,而非对时间性更深层次的理解。
  • 逐点误差未能捕捉如趋势保持和状态段鲁棒性等结构性特征。
  • 三维评估视角(统计保真性、结构一致性、决策相关性)能提供更丰富的预测评估。
  • 窗口级诊断揭示被全局平均值掩盖的变动性,并强调在特定状态下的表现。
  • 不存在普遍的冠军;应以领域特定目标来指导模型评估与部署。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。