[论文解读] Automated threshold selection and associated inference uncertainty for univariate extremes
本文提出了一种基于广义帕累托分布(GPD)的单变量极端值分析自动化阈值选择方法,直接解决阈值选择中的偏差-方差权衡问题。该方法提出了一种新颖的程序,以量化并传播阈值选择带来的不确定性至高分位数推断,通过模拟数据和River Nidd洪水数据集验证,其在准确性与鲁棒性方面优于现有方法。
Threshold selection is a fundamental problem in any threshold-based extreme value analysis. While models are asymptotically motivated, selecting an appropriate threshold for finite samples is difficult and highly subjective through standard methods. Inference for high quantiles can also be highly sensitive to the choice of threshold. Too low a threshold choice leads to bias in the fit of the extreme value model, while too high a choice leads to unnecessary additional uncertainty in the estimation of model parameters. We develop a novel methodology for automated threshold selection that directly tackles this bias-variance trade-off. We also develop a method to account for the uncertainty in the threshold estimation and propagate this uncertainty through to high quantile inference. Through a simulation study, we demonstrate the effectiveness of our method for threshold selection and subsequent extreme quantile estimation, relative to the leading existing methods, and show how the method's effectiveness is not sensitive to the tuning parameters. We apply our method to the well-known, troublesome example of the River Nidd dataset.
研究动机与目标
- 解决极端值分析中阈值选择的关键挑战,其中低阈值引入偏差,高阈值增加方差。
- 开发一种客观、自动化的阈值选择方法,以平衡偏差与方差,减少传统视觉或启发式方法中的主观性。
- 量化并传播从阈值选择到重现期水平推断的不确定性,提升极端分位数估计的可靠性。
- 通过模拟研究和真实水文数据,证明该方法在性能上优于现有自动化与主观阈值选择技术。
- 提供一种可扩展至协变量依赖模型的框架,拓展其在独立同分布数据之外的应用范围。
提出的方法
- 提出一种基于最小化偏差-方差权衡准则的新型自动化阈值选择程序,采用基于似然的优化框架。
- 通过将阈值选择建模为随机变量,并利用自助法技术传播该不确定性,实现其在推断中的量化。
- 以参数稳定性图为基础,但通过正式推断增强,识别出形状参数 ξ 在更高阈值下保持稳定的最低阈值。
- 应用基于自助法的重抽样方法,生成多个阈值与参数估计,实现对重现期预测的完整不确定性量化。
- 利用形状参数 ξ 和尺度参数 σᵤ 的GPD模型对选定阈值以上的超额值进行建模,确保渐近有效性。
- 通过在自助样本中聚合结果,将阈值不确定性整合到重现期推断中,生成更可靠的预测区间。
实验结果
研究问题
- RQ1如何在单变量极端值分析中实现阈值选择的自动化,以最小化偏差-方差权衡?
- RQ2阈值不确定性对极端分位数估计精度的影响是什么,如何对其进行恰当量化?
- RQ3与现有自动化及主观阈值选择技术相比,所提方法在偏差、方差和均方误差方面的表现如何?
- RQ4该方法能否可靠处理极端观测值有限的挑战性数据集,如River Nidd洪水数据?
- RQ5与单阈值方法相比,传播阈值不确定性在多大程度上提升了重现期推断的可靠性?
主要发现
- 在所有模拟情景中,所提方法在分位数估计中实现了最低的均方根误差(RMSE),尤其在偏差与方差方面显著优于Danielsson等(2001)和Danielsson等(2019)的方法。
- 对于样本量 n=2000 的正态分布数据,EQD方法在 0.95 分位数处实现了最低的 RMSE(0.214)与方差(0.038),优于Northrop与Wadsworth的方法。
- 在更大样本量(n=20,000)下,EQD方法保持最低方差并具备有竞争力的偏差,从而在所有分位数水平下均取得最小的 RMSE。
- 该方法在 n=2000 和 n=20,000 时分别选择中位数阈值为数据的 75% 和 87.5%,表明其阈值选择稳定且基于数据驱动。
- 在River Nidd数据集中,该方法提供了更可靠的重现期估计,预测区间更窄,展示了在真实世界应用中的鲁棒性。
- 基于自助法的不确定性传播成功捕捉了阈值选择的变异性,从而为极端分位数推断提供了更现实、更可靠的结论。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。