[Paper Review] Lead-lag detection and network clustering for multivariate time series with an application to the US equity market
This paper proposes an unsupervised method for detecting lead-lag clusters in multivariate time series by modeling pairwise lead-lag relationships as a directed weighted network and applying high-imbalance clustering. Applied to US equity data, the method identifies statistically significant clusters that generate a predictive trading signal with a 10% annualized volatility and a Sharpe ratio of 0.62, outperforming the S&P 500 and showing low market correlation.
In multivariate time series systems, it has been observed that certain groups of variables partially lead the evolution of the system, while other variables follow this evolution with a time delay; the result is a lead-lag structure amongst the time series variables. In this paper, we propose a method for the detection of lead-lag clusters of time series in multivariate systems. We demonstrate that the web of pairwise lead-lag relationships between time series can be helpfully construed as a directed network, for which there exist suitable algorithms for the detection of pairs of lead-lag clusters with high pairwise imbalance. Within our framework, we consider a number of choices for the pairwise lead-lag metric and directed network clustering components. Our framework is validated on both a synthetic generative model for multivariate lead-lag time series systems and daily real-world US equity prices data. We showcase that our method is able to detect statistically significant lead-lag clusters in the US equity market. We study the nature of these clusters in the context of the empirical finance literature on lead-lag relations and demonstrate how these can be used for the construction of predictive financial signals.
Motivation & Objective
- To detect lead-lag clusters in high-dimensional multivariate time series using unsupervised learning.
- To model pairwise lead-lag relationships as a directed weighted network for systematic analysis.
- To identify statistically significant clusters that can serve as predictive signals in financial markets.
- To validate the method on synthetic data and real-world US equity price data.
- To demonstrate the utility of lead-lag clustering for constructing robust, low-correlation trading signals.
Proposed method
- Constructs a directed weighted network where nodes represent time series and edges represent lead-lag relationships based on a chosen pairwise metric.
- Uses a state-of-the-art directed network clustering algorithm to detect communities with high cut imbalance, indicating strong directional influence between clusters.
- Employs multiple pairwise lead-lag metrics (e.g., cross-correlation with time lag, transfer entropy) and evaluates their performance via synthetic experiments.
- Applies volatility normalization and rolling optimization to improve signal stationarity and reliability in financial forecasting.
- Generates a predictive trading signal by ranking stocks within clusters based on lagged returns and rebalancing weekly.
- Validates signal performance using a Monte Carlo ablation study with permuted cluster labels to test statistical significance.
Experimental results
Research questions
- RQ1Can lead-lag relationships in multivariate time series be effectively modeled as a directed network to uncover latent cluster structures?
- RQ2How do different pairwise lead-lag metrics compare in capturing meaningful temporal dependencies in high-dimensional systems?
- RQ3Are the detected lead-lag clusters in US equity data statistically significant and not explainable by existing financial hypotheses?
- RQ4Can lead-lag cluster structures be leveraged to generate a predictive, low-correlation trading signal in noisy financial markets?
- RQ5How robust is the predictive signal under permutation of cluster labels, indicating the necessity of the cluster structure?
Key findings
- The method successfully detects statistically significant lead-lag clusters in US equity returns that are not explained by three prominent existing lead-lag hypotheses in empirical finance.
- The constructed trading signal achieves an annualized Sharpe ratio of 0.62 with a one-sided p-value < 0.004, significantly outperforming the S&P 500’s 0.40 Sharpe ratio.
- The signal exhibits a low correlation (0.04) with the market return, indicating it is not a market-wide exposure and likely captures a unique predictive signal.
- The performance of the signal decays after 2012, coinciding with a drop in clustering persistence and increased market informational efficiency.
- Ablation testing shows the probability of observing a Sharpe ratio ≥ 0.62 under the null hypothesis (random clusters) is p < 0.005, confirming the significance of the lead-lag structure.
- The signal’s mean daily return is 2.4 basis points, slightly below the market’s 3.0 basis points, but with improved risk-adjusted performance and low turnover.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.