Skip to main content
QUICK REVIEW

[论文解读] New g%AIC, g%AICc, g%BIC, and Power Divergence Fit Statistics Expose Mating between Modern Humans, Neanderthals and other Archaics

Peter J. Waddell, Xi Li Tan|arXiv (Cornell University)|Dec 31, 2012
Pleistocene-Era Hominins and Archaeology参考文献 25被引用 3
一句话总结

本文引入了新的信息准则——g%AIC、g%AICc、g%BIC,以及基于幂发散的拟合统计量,通过将信息理论原则与g%SD准则相结合,改进了遗传数据的模型选择。应用于智人和古人类基因组时,这些方法为古代杂交提供了强有力的统计证据,包括直立人对丹尼索瓦人的显著基因贡献。

ABSTRACT

The purpose of this article is to look at how information criteria, such as AIC and BIC, relate to the g%SD fit criterion derived in Waddell et al. (2007, 2010a). The g%SD criterion measures the fit of data to model based on a normalized weighted root mean square percentage deviation between the observed data and model estimates of the data, with g%SD = 0 being a perfectly fitting model. However, this criterion may not be adjusting for the number of parameters in the model comprehensively. Thus, its relationship to more traditional measures for maximizing useful information in a model, including AIC and BIC, are examined. This results in an extended set of fit criteria including g%AIC and g%BIC. Further, a broader range of asymptotically most powerful fit criteria of the power divergence family, which includes maximum likelihood (or minimum G^2) and minimum X^2 modeling as special cases, are used to replace the sum of squares fit criterion within the g%SD criterion. Results are illustrated with a set of genetic distances looking particularly at a range of Jewish populations, plus a genomic data set that looks at how Neanderthals and Denisovans are related to each other and modern humans. Evidence that Homo erectus may have left a significant fraction of its genome within the Denisovan is shown to persist with the new modeling criteria.

研究动机与目标

  • 通过引入类似AIC和BIC的模型复杂度调整,解决g%SD准则的局限性。
  • 通过将信息准则扩展至幂发散族,构建更全面的群体基因组学模型选择框架。
  • 改进对现代人类、尼安德特人、丹尼索瓦人及其他古人类之间古老混合事件的检测能力。
  • 利用增强的统计拟合准则,评估直立人是否对丹尼索瓦人谱系有显著的基因贡献。
  • 为遗传距离分析提供一种稳健的信息理论替代方法,替代传统的最小二乘法拟合。

提出的方法

  • 通过用幂发散统计量(包括最小G²(基于似然)和最小卡方X²(皮尔逊卡方)作为特例)替代原始g%SD准则中的平方偏差和,改进g%SD准则。
  • 通过将幂发散拟合嵌入信息理论框架,推导出新的信息准则——g%AIC、g%AICc(小样本校正)和g%BIC。
  • 将新拟合统计量应用于不同人类群体(包括犹太群体)的遗传距离矩阵,以及尼安德特人、丹尼索瓦人和现代人类的全基因组数据。
  • 使用幂发散族中渐近最有效的统计量评估模型拟合度,同时对参数数量施加惩罚。
  • 通过信息准则进行模型比较,以对混合与分化事件的进化模型进行排序。
  • 通过比较多个分化参数下的模型拟合度,并评估不同数据集间的一致性,验证结果。

实验结果

研究问题

  • RQ1与传统的g%SD相比,新的g%AIC和g%BIC准则是否能提升群体基因组学中模型选择的准确性?
  • RQ2使用新的拟合统计量,现代人类、尼安德特人和丹尼索瓦人之间杂交的统计证据是什么?
  • RQ3根据新的模型选择框架,直立人是否对丹尼索瓦人谱系有显著的基因贡献?
  • RQ4基于幂发散的拟合统计量在检测古人类混合方面,与经典最小二乘法和似然法相比表现如何?
  • RQ5新准则在保持对遗传分化模式敏感性的同时,多大程度上考虑了模型复杂度?

主要发现

  • 新的g%AIC、g%AICc和g%BIC准则相比原始g%SD准则,提供了更稳健、更全面的模型选择框架。
  • 幂发散族的拟合统计量成功替代了g%SD框架中的平方和,使模型更契合基于似然的推断。
  • 使用新的信息准则,对现代人类、尼安德特人和丹尼索瓦人之间杂交的统计支持非常有力。
  • 即使在应用改进的模型选择准则后,直立人对丹尼索瓦人谱系的基因贡献证据依然显著。
  • 新拟合统计量比以往方法更能可靠地检测到细微的遗传分化模式,尤其在复杂的混合情景中。
  • 结果表明,扩展的信息准则在识别古人类基因组最佳拟合进化模型方面非常有效。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。