[논문 리뷰] Are We Winning the Wrong Game? Revisiting Evaluation Practices for Long-Term Time Series Forecasting
이 논문은 장기 시계열 예측 평가가 합계 지향 지표(MSE/MAE)에 과도하게 의존하는 경향이 있으며, 통계적 충실도, 구조적 일관성, 의사결정 관련성에 초점을 맞춘 다차원 진단적 평가 프레임워크를 제안하여 리더보드 추종에서 맥락 인식 예측으로 전환한다.
Long-term time series forecasting (LTSF) is widely recognized as a central challenge in data mining and machine learning. LTSF has increasingly evolved into a benchmark-driven ''GAME,'' where models are ranked, compared, and declared state-of-the-art based primarily on marginal reductions in aggregated pointwise error metrics such as MSE and MAE. Across a small set of canonical datasets and fixed forecasting horizons, progress is communicated through leaderboard-style tables in which lower numerical scores define success. In this GAME, what is measured becomes what is optimized, and incremental error reduction becomes the dominant currency of advancement. We argue that this metric-centric regime is not merely incomplete, but structurally misaligned with the broader objectives of forecasting. In real-world settings, forecasting often prioritizes preserving temporal structure, trend stability, seasonal coherence, robustness to regime shifts, and supporting downstream decision processes. Optimizing aggregate pointwise error does not necessarily imply modeling these structural properties. As a result, leaderboard improvement may increasingly reflect specialization in benchmark configurations rather than a deeper understanding of temporal dynamics. This paper revisits LTSF evaluation as a foundational question in data science: what does it mean to measure forecasting progress? We propose a multi-dimensional evaluation perspective that integrates statistical fidelity, structural coherence, and decision-level relevance. By challenging the current metric monoculture, we aim to redirect attention from winning benchmark tables toward advancing meaningful, context-aware forecasting.
연구 동기 및 목표
- 집계된 점오차에 초점을 맞춘 현행 LTSF 평가 관행의 적절성에 의문을 제기한다.
- 통계적 충실도, 구조적 일관성, 의사결정 수준의 관련성을 포함하는 다차원 평가 프레임워크를 옹호한다.
- 진단적 보고 및 창(window) 수준 분석을 장려하여 구조적 조건 전반에 걸친 성능을 드러낸다.
- 보편적 리더십 지배보다는 도메인별 맥락 인식 평가를 촉진한다.
제안 방법
- LTSF의 벤치마크 주도 평가 체계와 그 인센티브를 비판적으로 분석한다.
- 통계적 충실도, 구조적 일관성, 의사결정 관련성의 3차원 평가 프레임워크를 제안한다.
- 진단적 보고와 윈도우-레벨 분석을 제시하여 서로 다른 체제에서의 성능 변동성을 평가한다.
- 보편적 리더를 지배하기보다는 도메인별 목표와 조건부 성능을 강조하는 보고를 옹호한다.
- 다양한 평가 관행을 설명하기 위해 emerging benchmarking initiatives(예: TSFM-Bench, TIME, TFB)를 참조하고 구축한다.
실험 결과
연구 질문
- RQ1현재 LTSF의 평가 체계가 한계적인 MSE/MAE 감소를 넘어 진짜 예측 목표와 일치하는가?
- RQ2LTSF 평가를 통계적 충실도, 시간적 구조, 의사결정 유용성의 균형을 맞추도록 어떻게 재구성할 수 있는가?
- RQ3다양한 구조적 조건에서 모델 행동을 드러낼 수 있는 창문 수준 진단과 같은 실제 평가 관행은 무엇인가?
- RQ4데이터 세트와 수평선에 상관없이 보편적인 챔피언이 존재하는가, 아니면 평가를 도메인 목표에 맥락화해야 하는가?
주요 결과
- 리더보드 중심의 진척은 종종 벤치마크별 적응을 반영하며 더 깊은 시간적 이해를 반영하지 않는다.
- 점오차는 추세 보존과 구간 변화에 대한 강건성 같은 구조적 특성을 포착하지 못한다.
- 통계적 충실도, 구조적 일관성, 의사결정 관련성의 3차원 평가 시야가 더 풍부한 예측 평가를 제공한다.
- 윈도우 수준 진단은 글로벌 평균으로 가려진 변동성을 드러내고 체제별 성능을 강조한다.
- 보편적 챔피언은 존재하지 않으며 도메인별 목표가 모델 평가 및 배치를 안내해야 한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.