[논문 리뷰] Integrating Inductive Biases in Transformers via Distillation for Financial Time Series Forecasting
다양한 귀납 편향(인과성, 국소성, 주기성)을 하나의 Transformer로 합성하는 증류 프레임워크 TIPS를 소개하여, 체제 변화 하의 금융 시계열 예측을 개선하고 더 낮은 추론 비용으로 최첨단 결과를 달성한다.
Transformer-based models have been widely adopted for time-series forecasting due to their high representational capacity and architectural flexibility. However, many Transformer variants implicitly assume stationarity and stable temporal dynamics -- assumptions routinely violated in financial markets characterized by regime shifts and non-stationarity. Empirically, state-of-the-art time-series Transformers often underperform even vanilla Transformers on financial tasks, while simpler architectures with distinct inductive biases, such as CNNs and RNNs, can achieve stronger performance with substantially lower complexity. At the same time, no single inductive bias dominates across markets or regimes, suggesting that robust financial forecasting requires integrating complementary temporal priors. We propose TIPS (Transformer with Inductive Prior Synthesis), a knowledge distillation framework that synthesizes diverse inductive biases -- causality, locality, and periodicity -- within a unified Transformer. TIPS trains bias-specialized Transformer teachers via attention masking, then distills their knowledge into a single student model with regime-dependent alignment across inductive biases. Across four major equity markets, TIPS achieves state-of-the-art performance, outperforming strong ensemble baselines by 55%, 9%, and 16% in annual return, Sharpe ratio, and Calmar ratio, while requiring only 38% of the inference-time computation. Further analyses show that TIPS generates statistically significant excess returns beyond both vanilla Transformers and its teacher ensembles, and exhibits regime-dependent behavioral alignment with classical architectures during their profitable periods. These results highlight the importance of regime-dependent inductive bias utilization for robust generalization in non-stationary financial time series.
연구 동기 및 목표
- 체제 변화와 비정상성으로 인해 금융 시계열 예측에서 적응형 귀납 편향의 필요성을 제시한다.
- 순수한 다중 편향 결합이 편향 특화 모델이나 앙상블에 비해 성능을 저하시킨다는 것을 보여준다.
- 다양한 편향을 하나의 트랜스포머로 합성하는 증류 기반 프레임워크 TIPS를 제안하고 검증한다.
- TIPS가 주요 주식 시장에서 최첨단 성능을 달성하는 한편 추론 비용을 줄임을 보인다.
제안 방법
- 주된 선험(priors)을 반영하는 편향 특화 Transformer 교사들을 인코딩하는 주의 마스킹과 입력 설계를 통해 학습시킨다.
- 일곱 교사(여섯 편향 특화 교사와 일반 Transformer)로 Bias Teacher Ensemble을 구성하여 다양한 선험을 포착한다.
- 앙상블 예측을 단일 학생 Transformer로 증류하고, 경직된 모방을 피하기 위해 적극적 정규화를 사용한다.
- 온도 스케일링으로 소프트 앙상블 타깃을 구성하고 보정(calibration)을 개선하기 위해 레이블 스무딩을 적용한다.
- 무제한 주의(attention)를 갖춘 학생 모델을 학습시켜 사전을 합성하고, 강건성을 위해 Stochastic Weight Averaging을 사용한다.
- 체제 의존적 편향 활성화와 통계적 초과 수익을 보여주기 위한 분석을 제공한다.

실험 결과
연구 질문
- RQ1다양한 귀납 편향이 비정상적 금융 데이터에서 Transformer의 강인성을 향상시킬 수 있는가?
- RQ2다수의 편향을 순진하게 결합하는 것이 특화나 앙상블에 비해 성능을 저하시키는가?
- RQ3증류된 학생이 추론 효율성을 유지하면서 여러 선험을 효과적으로 합성할 수 있는가?
- RQ4편향 선험이 체제별로 활성화되어 수익성 있는 시장 상황과 일치하는가?
- RQ5기준 모델을 넘어 TIPS가 통계적으로 상당한 초과 수익을 제공하는 정도는 어느 정도인가?
주요 결과
- TIPS는 네 가지 주요 주식 시장 전반에서 가장 강한 성능을 보이며, 기준 대비 평균 샤프비율과 연간 수익률이 우수하다.
- 주의 마스킹을 통한 Bias Teacher Ensemble이 고전적 아키텍처의 앙상블과 일반 SOTA 모델을 능가하며, 아키텍처 이질성 없이 편향 인코딩의 효과를 보인다.
- 단일 학생으로의 증류가 편향 앙상블에 비해 상당한 이점을 제공하고 추론 시간을 약 7배 단축하여 단일 모델로도 앙상블 수준의 강건성을 가능하게 한다.
- 절단 실험(Ablation)에서 정규화 구성요소들(저온 증류, 레이블 스무딩, SWA)이 효과적인 편향 합성에 함께 필요함을 보여준다.
- 분석에 따르면 TIPS는 일반적인 Transformer를 넘는 통계적으로 유의미한 알파를 달성하여, 귀납 편향 합성으로부터 유익한 신호를 추출함을 시사한다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.