[论文解读] Goodness-of-fit testing based on a weighted bootstrap: A fast large-sample alternative to the parametric bootstrap
本文提出了一种快速的大样本替代方法,用于在参数模型中进行拟合优度检验,替代计算成本高昂的参数自展法,采用加权自展法(乘子自展法),避免从估计的模型中重新抽样。该方法在显著降低计算成本的同时,实现了与参数自展法相当的检验功效,尤其在高维或参数较多的情况下优势明显,并通过蒙特卡洛模拟研究和对三元金融数据的应用得到了验证。
The process comparing the empirical cumulative distribution function of the sample with a parametric estimate of the cumulative distribution function is known as the empirical process with estimated parameters and has been extensively employed in the literature for goodness-of-fit testing. The simplest way to carry out such goodness-of-fit tests, especially in a multivariate setting, is to use a parametric bootstrap. Although very easy to implement, the parametric bootstrap can become very computationally expensive as the sample size, the number of parameters, or the dimension of the data increase. An alternative resampling technique based on a fast weighted bootstrap is proposed in this paper, and is studied both theoretically and empirically. The outcome of this work is a generic and computationally efficient multiplier goodness-of-fit procedure that can be used as a large-sample alternative to the parametric bootstrap. In order to approximately determine how large the sample size needs to be for the parametric and weighted bootstraps to have roughly equivalent powers, extensive Monte Carlo experiments are carried out in dimension one, two and three, and for models containing up to nine parameters. The computational gains resulting from the use of the proposed multiplier goodness-of-fit procedure are illustrated on trivariate financial data. A by-product of this work is a fast large-sample goodness-of-fit procedure for the bivariate and trivariate t distribution whose degrees of freedom are fixed.
研究动机与目标
- 解决当样本量、参数数量或维度增加时,多变量拟合优度检验中参数自展法计算成本过高的问题。
- 提出一种计算高效的、基于大样本的参数自展法替代方法,采用加权自展法(乘子自展法)实现。
- 评估在不同样本量、维度(1–3)以及最多九个参数的模型下,加权自展法与参数自展法在检验功效上的等价性。
- 为固定自由度的多变量t分布提供一种快速、通用的拟合优度检验程序。
提出的方法
- 基于经验过程在参数估计下的渐近线性化,提出一种乘子自展程序,用从高斯过程重新抽样替代从模型中重新抽样。
- 利用乘子中心极限定理近似原假设下检验统计量的分布,避免对参数模型进行重复模拟。
- 推导多变量t分布的CDF和PDF关于位置、尺度和相关性参数的显式梯度表达式,以实现原假设下检验统计量的高效计算。
- 通过生成独立同分布的标准正态变量并利用估计的影响函数进行加权,实现加权自展,避免从模型分布中进行昂贵的模拟。
- 利用生成的自展重抽样结果,通过经验分位数计算p值,其计算复杂度与模型模拟无关。
- 将该方法应用于三元金融数据,以展示计算效率的提升和实际可行性。
实验结果
研究问题
- RQ1样本量需多大时,加权自展法才能实现与参数自展法相当的检验功效?
- RQ2加权自展法能否作为高维或高参数模型中参数自展法的计算高效替代方法?
- RQ3在多变量设置下,加权自展法相比参数自展法的计算效率提升程度如何?
- RQ4加权自展法在固定自由度的多变量t分布中是否能保持准确的大小和功效?
- RQ5在不同维度和参数数量下,加权自展法与参数自展法在经验第一类错误率和功效方面的表现如何比较?
主要发现
- 当样本量中等偏大时,加权自展法可实现与参数自展法相当的检验功效,仿真结果在维度1至3及最多九参数的模型中均显示两者等价。
- 广泛的蒙特卡洛实验表明,即使在高维设置下参数自展法变得不可行,加权自展法仍能保持准确的经验大小和功效。
- 该方法显著降低了计算成本:参数自展法需为每个自展重抽样从模型中进行N次模拟,而加权自展法通过使用独立同分布的正态变量和影响函数避免了这一过程。
- 对于三元金融数据,所提方法相比参数自展法将计算时间减少了几个数量级,实现了实际应用的可行性。
- 作为副产品,开发了一种针对固定自由度的二元和三元t分布的快速、大样本拟合优度检验程序,利用CDF和PDF的解析梯度。
- 理论依据来自乘子中心极限定理,表明加权自展法渐近地复制了带参数估计的经验过程的原假设分布。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。