Skip to main content
QUICK REVIEW

[论文解读] Kernel Approximate Bayesian Computation for Population Genetic Inferences

Shigeki Nakagome, Kenji Fukumizu|arXiv (Cornell University)|May 15, 2012
Markov Chains and Monte Carlo Methods参考文献 26被引用 13
一句话总结

本文提出基于核方法的近似贝叶斯计算(核ABC)用于群体遗传推断,克服了传统ABC方法的局限性,通过使用核方法处理大量汇总统计量,同时保持性能。与标准ABC不同,后者因汇总统计量不足和容忍度非零而导致接受率低及近似误差,核ABC利用概率测度的核嵌入及特征空间中的差异最小化,在处理复杂高维数据时仍能实现高精度推断。

ABSTRACT

Approximate Bayesian computation (ABC) is a likelihood-free approach for Bayesian inferences based on a rejection algorithm method that applies a tolerance of dissimilarity between summary statistics from observed and simulated data. Although several improvements to the algorithm have been proposed, none of these improvements avoid the following two sources of approximation: 1) lack of sufficient statistics: sampling is not from the true posterior density given data but from an approximate posterior density given summary statistics; and 2) non-zero tolerance: sampling from the posterior density given summary statistics is achieved only in the limit of zero tolerance. The first source of approximation can be improved by adding a summary statistic, but an increase in the number of summary statistics could introduce additional variance caused by the low acceptance rate. Consequently, many researchers have attempted to develop techniques to choose informative summary statistics. The present study evaluated the utility of a kernel-based ABC method (Fukumizu et al. 2010, arXiv:1009.5736 and 2011, NIPS 24: 1549-1557) for complex problems that demand many summary statistics. Specifically, kernel ABC was applied to population genetic inference. We demonstrate that, in contrast to conventional ABCs, kernel ABC can incorporate a large number of summary statistics while maintaining high performance of the inference.

研究动机与目标

  • 解决传统近似贝叶斯计算(ABC)在群体遗传学中依赖不足汇总统计量和容忍度非零的局限性。
  • 克服增加汇总统计量以提升推断性能与因高维数据导致接受率下降之间的权衡。
  • 评估基于核的ABC在需要大量汇总统计量的复杂群体遗传推断问题中的性能。
  • 证明核ABC即使在使用大量汇总统计量时仍能保持高推断精度,优于标准ABC方法。
  • 为具有复杂结构的现代群体遗传数据集提供一种可扩展且统计上可靠的传统ABC替代方案。

提出的方法

  • 利用核方法将观测数据与模拟数据的概率测度嵌入再生核希尔伯特空间(RKHS),实现分布的非参数比较。
  • 采用最大均值差异(MMD)作为观测数据与模拟数据汇总统计量经验分布之间的差异度量。
  • 用基于核的相似性度量替代传统ABC中的容忍度排斥采样,实现更平滑、更高效的后验分布采样。
  • 将核ABC应用于群体遗传模型,使用从基因数据中提取的广泛汇总统计量(如等位基因频率、FST、Tajima's D)。
  • 在RKHS中采用核密度估计方法近似后验分布,无需显式似然函数。
  • 通过交叉验证优化核带宽及汇总统计量选择,以提高后验近似精度。

实验结果

研究问题

  • RQ1核ABC能否在不显著降低接受率的情况下,有效处理群体遗传推断中的大量汇总统计量?
  • RQ2在使用高维汇总统计量时,核ABC与传统ABC在后验估计精度方面有何差异?
  • RQ3核ABC在多大程度上减少了ABC中因汇总统计量不足和容忍度非零导致的近似误差?
  • RQ4在真实群体遗传模型中,核ABC对核函数和带宽选择是否具有鲁棒性?
  • RQ5核ABC能否在具有复杂人口历史的复杂群体遗传数据集中实际应用?

主要发现

  • 即使使用大量汇总统计量,核ABC仍能保持高推断精度,而标准ABC通常因接受率过低导致性能下降。
  • 该方法通过有效利用高维、信息丰富的统计量,显著减少了因汇总统计量不足导致的近似误差。
  • 实证结果表明,核ABC在多种群体遗传模型(包括隔离-迁移模型和种群瓶颈情景)中均优于传统ABC,后验估计更优。
  • 在RKHS中使用MMD作为差异度量,相比标准ABC中的容忍度排斥采样,实现了更稳定、更高效的采样。
  • 核ABC对核函数选择和带宽选择具有鲁棒性,实际应用中交叉验证可进一步提升性能。
  • 该方法即使在高维汇总统计量空间中,也能准确推断复杂人口参数(如迁移率和分化时间)。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。