[论文解读] Information Criteria for Deciding between Normal Regression Models
本文澄清了在物理科学中选择正态回归模型时,信息准则(特别是AIC和AICc)的正确应用。研究指出,AICc仅在误差方差未知时适用;当误差条已知时,应使用AIC。主要贡献是提出了在已知误差方差条件下对AIC差异的显著性检验,该方法可替代Akaike权重,并提高了天体物理学和物理数据分析中模型选择的可靠性。
Regression models fitted to data can be assessed on their goodness of fit, though models with many parameters should be disfavored to prevent over-fitting. Statisticians' tools for this are little known to physical scientists. These include the Akaike Information Criterion (AIC), a penalized goodness-of-fit statistic, and the AICc, a variant including a small-sample correction. They entered the physical sciences through being used by astrophysicists to compare cosmological models; e.g., predictions of the distance-redshift relation. The AICc is shown to have been misapplied, being applicable only if error variances are unknown. If error bars accompany the data, the AIC should be used instead. Erroneous applications of the AICc are listed in an appendix. It is also shown how the variability of the AIC difference between models with a known error variance can be estimated. This yields a significance test that can potentially replace the use of `Akaike weights' for deciding between such models. Additionally, the effects of model misspecification are examined. For regression models fitted to data sets without (rather than with) error bars, they are major: the AICc may be shifted by an unknown amount. The extent of this in the fitting of physical models remains to be studied.
研究动机与目标
- 解决物理科学中AICc的广泛误用问题,特别是在天体物理学中,误差条通常已提供。
- 澄清AICc仅在误差方差未知时有效;当误差条已知时,应使用AIC。
- 在已知误差方差条件下,开发AIC差异的显著性检验,从而以经典假设检验替代Akaike权重。
- 研究模型误设对AIC和AICc性能的影响,特别是在缺乏误差条的情况下。
- 识别并列出近期天体物理学文献中AICc的错误使用案例,强调方法论修正的必要性。
提出的方法
- 在一般框架下推导信息准则,将其与Kullback–Leibler散度及基于最大似然估计(MLE)的估计方法联系起来。
- 将该框架应用于具有已知或未知误差方差的线性正态回归模型,明确区分AIC与AICc的适用条件。
- 使用非中心卡方分布来建模在模型误设条件下AIC差异的抽样变异性。
- 通过估计ΔAIC的标准误,提出AIC差异的假设检验,从而实现p值计算。
- 采用自助法程序(如Tan & Biswas, 2012所述)对AICc的变异性进行经验验证,但本研究的重点在于理论依据。
- 分析AICc在模型误设下的渐近行为,表明校正项可能被偏差所淹没。
实验结果
研究问题
- RQ1在误差条已知的情况下,何时应使用AIC而非AICc进行回归模型选择?
- RQ2在已知误差方差的两个模型之间,AIC值的差异能否用于执行经典显著性检验?
- RQ3模型误设如何影响AICc在模型选择中的可靠性,特别是在误差方差未知时?
- RQ4错误地将AICc应用于具有明确误差条的数据会产生什么后果?该错误在天体物理学中有多普遍?
- RQ5AIC差异的变异性在多大程度上可被估计,从而可替代Akaike权重在模型比较中的使用?
主要发现
- 当误差方差已知或以误差条形式提供时,不应使用AICc;此时AIC是正确选择。
- 当误差方差已知时,AIC差异的抽样变异性可被估计,从而可进行具有p值的正式显著性检验。
- 所提出的基于AIC差异变异性显著性检验可替代Akaike权重,后者缺乏坚实的频率学派基础。
- AICc的错误应用在天体物理学文献中普遍存在,包括Biesiada & Piórkowska (2009)、February et al. (2010)等论文,详见附录B。
- 在模型误设下,AICc的校正项可能被偏差所淹没,尤其在大样本中,从而降低其有效性。
- 对于无误差条的数据,由于误设的影响,AICc可能被未知量偏移,从而削弱其在模型选择中的可靠性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。