[论文解读] Evaluation of binary classifiers for asymptotically dependent and independent extremes
本文提出了一种新颖的风险函数,用于在极端值场景下评估二元分类器,解决了罕见事件预测中的数据不平衡问题。通过基于事件发生概率重新加权误分类成本,该方法确保了分类器之间的公平比较,并在多变量正则变体框架下实现了理论一致性与实证验证,应用于河流流量预测。
Machine learning classification methods usually assume that all possible classes are sufficiently present within the training set. Due to their inherent rarities, extreme events are always under-represented and classifiers tailored for predicting extremes need to be carefully designed to handle this under-representation. In this paper, we address the question of how to assess and compare classifiers with respect to their capacity to capture extreme occurrences. This is also related to the topic of scoring rules used in forecasting literature. In this context, we propose and study a risk function adapted to extremal classifiers. The inferential properties of our empirical risk estimator are derived under the framework of multivariate regular variation and hidden regular variation. A simulation study compares different classifiers and indicates their performance with respect to our risk function. To conclude, we apply our framework to the analysis of extreme river discharges in the Danube river basin. The application compares different predictive algorithms and test their capacity at forecasting river discharges from other river stations.
研究动机与目标
- 解决在标准风险函数倾向于选择‘始终为负’预测的罕见极端事件分类问题评估挑战。
- 开发一种风险函数,公平惩罚过度悲观和过度乐观的分类器,避免对平凡解的偏倚。
- 在多变量正则变体和隐性正则变体框架下,建立风险函数的理论一致性。
- 通过模拟和多瑙河流域极端河流流量的实际应用,对方法进行实证验证。
- 利用在新风险函数下优化的线性分类器,识别极端行为的关键预测因子。
提出的方法
- 提出加权损失函数 $ l_u(g;(m{x},y)) = \frac{1}{\mathbb{P}(Y^{(u)}=1 \text{ 或 } g(\bm{X};u)=1)} \mathds{1}\{g(\bm{x};u)\neq y\} $,通过事件与预测概率的并集对误分类成本进行归一化。
- 定义相关风险 $ R^{(u)}(g) = \mathbb{E}[l_u(g;(m{X},Y^{(u)}))] $,其取值范围为 [0,1],且‘始终为负’和‘始终为正’分类器的风险值均为 1。
- 在多变量正则变体和隐性正则变体框架下,建立经验风险估计器的渐近理论,证明其弱收敛于高斯过程。
- 利用经验过程理论和 bracketing entropy,基于适当的矩条件和连续性条件,验证风险估计器的紧致性与弱收敛性。
- 将该框架应用于线性分类器 $ b_{\bm{\theta}}(\bm{x},h) = \mathds{1}\{\bm{\theta}^T\bm{x}>1 \text{ 或 } h>1\} $,表明最小化风险可一致估计极端依赖结构。
- 通过函数 delta 方法和 Lindeberg 型条件,推导经验风险估计器的渐近正态性,适用于随机过程。
实验结果
研究问题
- RQ1当正类样本极为稀少时,如何构建一个公平的风险函数以评估极端事件的二元分类器?
- RQ2在多变量正则变体和隐性正则变体框架下,经验风险估计器的渐近性质是什么?
- RQ3所提出的风险函数能否在高维设置下一致识别出极端行为的最关键预测因子?
- RQ4与经典的临界成功指数等指标相比,新风险函数在极端事件预测中的表现如何?
- RQ5该框架在如河流流量预测等实际应用中,能在多大程度上提升预测性能?
主要发现
- 所提出的风险函数 $ R^{(u)}(g) $ 确保了‘始终为负’和‘始终为正’分类器的风险值均为 1,为公平比较提供了明确基准。
- 在多变量正则变体下,经验风险估计器弱收敛于高斯过程,收敛速度取决于尾指数和维度。
- 对于线性分类器,最小化风险函数可一致估计极端依赖结构,识别出最具影响力的预测因子。
- 在多瑙河流域河流流量的应用中,该框架成功比较了多种预测算法,并识别出极端事件的关键贡献变量。
- 模拟研究证实,新风险函数能有效按分类器的真实极端预测能力进行排序,避免对朴素模型的偏倚。
- 理论分析表明,风险估计器渐近正态,收敛速度由尾概率 $ \mathbb{P}(H>u_n) \to 0 $ 决定,当 $ u_n \to \infty $ 时成立。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。