[论文解读] Extreme Value Analysis Without the Largest Values: What Can Be Done?
本文提出了一种函数型Hill估计量HEWE,通过将未观测到的极端值数量建模为一个参数,以处理重尾数据中缺失的极端值。该方法建立了对高斯过程的函数收敛性,并实现了尾指数、偏差和缺失极端值数量的联合估计,通过模拟和真实世界社交网络数据验证,在缺失情况下显著提升了参数估计性能。
In this paper we are concerned with the analysis of heavy-tailed data when a portion of the extreme values is unavailable. This research was motivated by an analysis of the degree distributions in a large social network. The degree distributions of such networks tend to have power law behavior in the tails. We focus on the Hill estimator, which plays a starring role in heavy-tailed modeling. The Hill estimator for this data exhibited a smooth and increasing "sample path" as a function of the number of upper order statistics used in constructing the estimator. This behavior became more apparent as we artificially removed more of the upper order statistics. Building on this observation we introduce a new version of the Hill estimator. It is a function of the number of the upper order statistics used in the estimation, but also depends on the number of unavailable extreme values. We establish functional convergence of the normalized Hill estimator to a Gaussian process. An estimation procedure is developed based on the limit theory to estimate the number of missing extremes and extreme value parameters including the tail index and the bias of Hill's estimator. We illustrate how this approach works in both simulations and real data examples.
研究动机与目标
- 解决当数据集中最大观测值缺失或未被观测到时,估计极端值参数的挑战。
- 开发一种统计框架,同时估计重尾分布中缺失极端值的数量和尾指数。
- 建模在缺失极端值情况下Hill估计量的行为,此时样本路径变得平滑且递增,与典型波动模式相反。
- 在适当的正则变异性假设下,为新估计量建立渐近理论,包括函数收敛至高斯过程。
- 提供实用的估计程序和计算工具,用于实际数据应用,包括社交网络度分布和自然灾害数据。
提出的方法
- 提出一种双参数函数型Hill估计量HEWE(θ;δ),其中θ为所用上端顺序统计量的比例,δ为缺失极端值的比例。
- 通过将归一化估计量建模为涉及布朗运动和由缺失极端值引起的偏差项的函数极限,推导出HEWE的渐近分布。
- 利用K"{o}mlos-Major-Tusn\'{a}dy逼近,将经验过程与布朗运动联系起来,从而推导出极限高斯过程。
- 建立归一化HEWE对具有特定协方差结构的高斯过程的函数收敛性,该结构取决于δ和θ。
- 推导出涉及正则变异性指数ρ和缩放函数A(n/k_n)的渐近偏差项II,该偏差项捕捉了缺失极端值对估计量的影响。
- 基于极限理论开发一种估计程序,从未观测数据中推断δ(缺失极端值数量)、γ(尾指数)和偏差参数。
实验结果
研究问题
- RQ1当最大值缺失,尤其是缺失极端值数量未知时,Hill估计量能否被调整以处理此类数据?
- RQ2当上端顺序统计量被系统性地移除时,Hill估计量的理论行为如何?
- RQ3如何在估计尾指数和偏差的同时,估计缺失极端值的数量?
- RQ4在缺失极端值条件下,函数型Hill估计量的渐近分布是什么?能否用于统计推断?
- RQ5所提出的方法能否应用于现实世界的重尾数据(如社交网络度分布),其中极端值可能未被观测到?
主要发现
- 函数型Hill估计量HEWE(θ;δ)以弱收敛方式趋于一个非退化的协方差结构依赖于θ和δ的高斯过程,从而支持统计推断。
- 在缺失极端值情况下,Hill估计量的偏差渐近正比于A(n/k_n)乘以δ和θ的函数,且在正则变异性下,缩放函数A(n/k_n)以O(1/√k_n)的速率衰减。
- 在模拟中,该方法成功估计了缺失极端值数量δ和尾指数γ,即使最多10%的最大值被移除,也能准确恢复。
- 在Google+社交网络数据中,观测到的平滑且递增的Hill图可归因于缺失的高阶度节点,模型估计约有1,000至2,000个极端入度未被观测到。
- R Shiny网络应用程序支持实时参数估计和交互式可视化,允许用户模拟缺失极端值,并比较缺失前后的估计性能。
- 在存在缺失极端值的情况下,该方法优于标准Hill估计,因为后者因缺少上尾而无法识别k选择的稳定区域。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。