[Paper Review] Modelling multivariate extreme value distributions via Markov trees
This paper proposes a tree-structured Markov random field model using bivariate extreme value distributions to approximate high-dimensional multivariate extreme value distributions, leveraging Prim’s algorithm with tail dependence coefficients to learn the tree structure. The method enables efficient inference on rare event probabilities, demonstrated effectively on upper Danube river discharge data with improved flooding probability estimates over empirical values.
Multivariate extreme value distributions are a common choice for modelling multivariate extremes. In high dimensions, however, the construction of flexible and parsimonious models is challenging. We propose to combine bivariate max-stable distributions into a Markov random field with respect to a tree. Although in general not max-stable itself, this Markov tree is attracted by a multivariate max-stable distribution. The latter serves as a tree-based approximation to an unknown max-stable distribution with the given bivariate distributions as margins. Given data, we learn an appropriate tree structure by Prim's algorithm with estimated pairwise upper tail dependence coefficients as edge weights. The distributions of pairs of connected variables can be fitted in various ways. The resulting tree-structured max-stable distribution allows for inference on rare event probabilities, as illustrated on river discharge data from the upper Danube basin.
Motivation & Objective
- To address the challenge of constructing flexible, parsimonious models for high-dimensional multivariate extreme value distributions.
- To develop a tractable approximation to an unknown multivariate extreme value distribution using a Markov tree structure.
- To enable efficient inference on rare event probabilities in complex, high-dimensional dependence structures.
- To integrate graphical models with extreme value theory by learning tree structures from data via tail dependence measures.
- To demonstrate the method’s effectiveness on real-world hydrological data from the upper Danube basin.
Proposed method
- Construct a Markov random field over a tree structure using bivariate extreme value distributions as conditional distributions.
- Learn the optimal tree structure using Prim’s algorithm with edge weights derived from estimated upper tail dependence coefficients or Kendall’s tau values.
- Fit bivariate Hüsler–Reiss distributions to connected pairs using moments estimator, M-estimator, or weighted least squares with k=65.
- Estimate marginal tail probabilities using the generalized Pareto distribution (GPD) fitted to excesses over high thresholds (p=0.9).
- Approximate the full multivariate extreme value distribution as the limit of the Markov tree, enabling inference on joint tail events.
- Use the resulting model to compute and compare probabilities of rare events, such as flooding at multiple stations.
Experimental results
Research questions
- RQ1Can a tree-structured Markov random field based on bivariate extreme value distributions effectively approximate a high-dimensional multivariate extreme value distribution?
- RQ2How well does the model based on estimated tail dependence coefficients (via Kendall’s tau or upper tail dependence) capture the true dependence structure of extreme events?
- RQ3To what extent does the tree-based model improve the estimation of rare event probabilities compared to empirical estimates?
- RQ4How do different fitting methods (moments, M-estimator, WLS) affect the accuracy of the resulting extreme value model?
- RQ5Can the method be effectively applied to real-world hydrological data to predict joint flooding events in river networks?
Key findings
- The tree-based model consistently overestimates empirical flooding probabilities, with estimates ranging from 9.26% to 9.72% for the 0.95 quantile at stations 4, 7, and 13, compared to the empirical value of 8.44%.
- For the 0.99 quantile, model estimates ranged from 2.53% to 2.65%, exceeding the empirical 1.88%, indicating improved detection of joint extreme events.
- The GPD-based marginal estimates showed higher tail probabilities (e.g., 7.02% for station 13 at 0.999 quantile), contributing to the overestimation of joint tail probabilities.
- The model based on Kendall’s tau (tree $ ilde{ au}^*$) yielded slightly lower estimates than the upper tail dependence-based tree ($ ilde{ au}^*$), suggesting sensitivity to dependence measure choice.
- The use of multiple fitting estimators (MM, M, WLS) produced nearly identical results, indicating robustness of the model to estimation method within the Hüsler–Reiss framework.
- The method successfully captures complex joint tail behavior in high-dimensional river discharge data, demonstrating practical utility for flood risk assessment.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.