Skip to main content
QUICK REVIEW

[논문 리뷰] LLM as a Risk Manager: LLM Semantic Filtering for Lead-Lag Trading in Prediction Markets

Sumin Kim, Minjae Kim|arXiv (Cornell University)|2026. 02. 04.
Stock Market Forecasting Methods인용 수 0
한 줄 요약

두 단계 프레임워크는 Granger 인과관계와 LLM 기반 시맨틱 필터를 결합하여 예측 시장의 선도–지연 관계를 순위화하고 Kalshi 데이터에서 거래 성과를 개선하며 하방 위험을 감소시킵니다.

ABSTRACT

Prediction markets provide a unique setting where event-level time series are directly tied to natural-language descriptions, yet discovering robust lead-lag relationships remains challenging due to spurious statistical correlations. We propose a hybrid two-stage causal screener to address this challenge: (i) a statistical stage that uses Granger causality to identify candidate leader-follower pairs from market-implied probability time series, and (ii) an LLM-based semantic stage that re-ranks these candidates by assessing whether the proposed direction admits a plausible economic transmission mechanism based on event descriptions. Because causal ground truth is unobserved, we evaluate the ranked pairs using a fixed, signal-triggered trading protocol that maps relationship quality into realized profit and loss (PnL). On Kalshi Economics markets, our hybrid approach consistently outperforms the statistical baseline. Across rolling evaluations, the win rate increases from 51.4% to 54.5%. Crucially, the average magnitude of losing trades decreases substantially from 649 USD to 347 USD. This reduction is driven by the LLM's ability to filter out statistically fragile links that are prone to large losses, rather than relying on rare gains. These improvements remain stable across different trading configurations, indicating that the gains are not driven by specific parameter choices. Overall, the results suggest that LLMs function as semantic risk managers on top of statistical discovery, prioritizing lead-lag relationships that generalize under changing market conditions.

연구 동기 및 목표

  • 이벤트 수준 예측 시장에서 강력한 선도–지연 관계를 식별하는 문제의 동기를 제시합니다.
  • 통계적 탐색과 시맨틱 검증을 결합하는 두 단계 프레임워크를 제안합니다.
  • LLM 기반 재순위가 Granger 기반 단독 스크리닝보다 거래 성능을 개선하는지 평가합니다.

제안 방법

  • 1단계는 로그-오즈 변환된 가격에 대한 Granger 인과관계를 사용하여 여러 시차(p∈{1,2,3,4,5})에 걸친 후보 선도–추종 쌍을 식별합니다.
  • 2단계는 LLM(GPT-5-nano)을 사용하여 이벤트 제목/설명을 바탕으로 이러한 후보를 재순위화하고, 그럴듯한 경제적 전달 메커니즘을 평가하여 타당성 점수를 부여합니다.
  • 거래 프로토콜: 각 선도–추종 쌍에 대해 선도 가격이 임계치를 초과하면 추종 거래를 시작하고, h일 보유하며, 샘플 밖 PnL을 측정합니다.
  • 최종 포트폴리오는 통계에서 상위 M개의 방향성 쌍(M=20)을 선택합니다. 비교를 위한 하이브리드 접근법도 포함합니다.
  • 평가는 Kalshi Economics 마켓에서 60일 학습 및 30일 테스트의 롤링 윈도우를 사용합니다.
Figure 1: Two-stage framework for leader–follower pair discovery in prediction markets. Stage 1 produces a candidate set of Top K directed pairs (K=100) ranked by Granger significance, and Stage 2 applies LLM-based semantic re-ranking to select the final Top M portfolio (M=20).
Figure 1: Two-stage framework for leader–follower pair discovery in prediction markets. Stage 1 produces a candidate set of Top K directed pairs (K=100) ranked by Granger significance, and Stage 2 applies LLM-based semantic re-ranking to select the final Top M portfolio (M=20).

실험 결과

연구 질문

  • RQ1LLM이 Granger 인과관계로 식별된 취약한 통계적 상관관계와 기계적으로 그럴듯한 선도–지연 관계를 구분할 수 있는가?
  • RQ2시맨틱 재순위가 Granger 기반 스크리닝을 넘어서 거래 성과를 개선하고 하방 위험을 감소시키는가?
  • RQ3다양한 보유 기간과 시장 상황에서 이익이 견고한가?

주요 결과

  • LLM 기반 시맨틱 필터링은 통계적 기준선에 비해 총 PnL에서 상당한 이익(+205%)을 창출합니다.
  • 하방 위험이 감소하여 평균 손실이 $649에서 $347로 감소합니다(감소율 46.5%).
  • 기본 설정에서 승률이 51.4%에서 54.5%로 상승합니다.
  • 같은 이벤트 및 서로 다른 이벤트 쌍에서도 손실 감소가 지속되며(40% 이상 개선).
  • 시맨틱 필터링은 선도 움직임이 큰 경우에 더 큰 이익을 제공합니다(5–10pt 및 10pt+ 움직임).
  • LLM이 선택한 쌍은 Granger Top-M 밖에서도 경제적으로 의미 있는 연계를 포함할 수 있으며, 통계적 순위 이상으로 질적 이득을 보여줍니다.
Figure 2: Signal-triggered trading protocol used to evaluate ranked lead-lag relationships from Figure 1 : leader price moves trigger follower trades, with direction determined by the Granger-induced trade sign and out-of-sample PnL used to evaluate the ranked pair list.
Figure 2: Signal-triggered trading protocol used to evaluate ranked lead-lag relationships from Figure 1 : leader price moves trigger follower trades, with direction determined by the Granger-induced trade sign and out-of-sample PnL used to evaluate the ranked pair list.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.