[Paper Review] Split Conformal Prediction and Non-Exchangeable Data
This paper extends split conformal prediction to dependent data by showing that coverage guarantees under exchangeability can be preserved for stationary β-mixing processes with only a small penalty. The method uses standard split conformal procedures but proves theoretical bounds on marginal and empirical coverage that match the order of iid results, validated empirically on time series data including financial and climate datasets.
Split conformal prediction (CP) is arguably the most popular CP method for uncertainty quantification, enjoying both academic interest and widespread deployment. However, the original theoretical analysis of split CP makes the crucial assumption of data exchangeability, which hinders many real-world applications. In this paper, we present a novel theoretical framework based on concentration inequalities and decoupling properties of the data, proving that split CP remains valid for many non-exchangeable processes by adding a small coverage penalty. Through experiments with both real and synthetic data, we show that our theoretical results translate to good empirical performance under non-exchangeability, e.g., for time series and spatiotemporal data. Compared to recent conformal algorithms designed to counter specific exchangeability violations, we show that split CP is competitive in terms of coverage and interval size, with the benefit of being extremely simple and orders of magnitude faster than alternatives.
Motivation & Objective
- Address the gap in conformal prediction theory for dependent data, particularly non-iid time series.
- Extend the theoretical guarantees of split conformal prediction—originally valid under data exchangeability—to stationary β-mixing processes.
- Demonstrate that standard split CP methods maintain strong coverage properties even under temporal dependence.
- Provide theoretical and empirical validation of marginal, empirical, and conditional coverage under dependence.
- Generalize the framework to non-stationary processes and other conformal prediction methods beyond split CP.
Proposed method
- Adapt the standard split conformal prediction framework by dividing data into training, calibration, and test sets.
- Train a conformity score function on the training set, using residuals or other user-defined functions.
- Compute the empirical (1−α)-quantile of conformity scores on the calibration set to define prediction intervals.
- Use the quantile to construct prediction sets for test points: C_{1−α}(X_i) = {y : s_train(X_i, y) ≤ q_{1−α}}.
- Theoretical analysis leverages β-mixing coefficients to bound coverage deviation from the i.i.d. case.
- Prove marginal and empirical coverage bounds that scale with the mixing rate and calibration set size, showing convergence to i.i.d. guarantees.
Experimental results
Research questions
- RQ1Can split conformal prediction maintain valid marginal coverage under temporal dependence, such as in β-mixing processes?
- RQ2How does the coverage error scale with the degree of dependence and calibration set size in dependent data?
- RQ3Can the same split CP framework achieve valid conditional coverage given covariate-specific events (e.g., high volatility)?
- RQ4To what extent do theoretical coverage bounds for dependent data match those under i.i.d. assumptions?
- RQ5Can the framework be extended beyond stationarity and applied to other conformal prediction methods like rank-one-out or risk-controlling sets?
Key findings
- For stationary β-mixing processes, split conformal prediction achieves marginal coverage with error O(1/n_cal) that matches the order of the i.i.d. case.
- Empirical coverage over test sets is consistently above the theoretical lower bound and converges to the nominal level as calibration set size increases.
- In financial time series (e.g., EURUSD, Brent oil, S&P 500), marginal coverage remains close to 90% nominal level (e.g., 89.19% with 1000 calibration points), improving with larger calibration sets.
- Conditional coverage is maintained across subtypes (e.g., uptrends, high volatility), with empirical rates near 90% even under conditioning, confirming robustness.
- Theoretical guarantees hold for non-stationary processes and extend to other conformal methods, such as rank-one-out and risk-controlling prediction sets.
- Experiments show that coverage remains stable and close to nominal levels even under high dependence, validating the small penalty introduced by dependence.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.