[Paper Review] AABC: approximate approximate Bayesian computation when simulating a large number of data sets is computationally infeasible
This paper introduces AABC (Approximate Approximate Bayesian Computation), a method that enables Bayesian inference for complex mechanistic models when simulating data is computationally prohibitive. By using a small number of simulations to build a non-mechanistic, resampled model, AABC allows efficient generation of large posterior samples, converging to the true ABC posterior as simulation size and data size increase.
Approximate Bayesian computation (ABC) methods perform inference on model-specific parameters of mechanistically motivated parametric statistical models when evaluating likelihoods is difficult. Central to the success of ABC methods is computationally inexpensive simulation of data sets from the parametric model of interest. However, when simulating data sets from a model is so computationally expensive that the posterior distribution of parameters cannot be adequately sampled by ABC, inference is not straightforward. We present approximate approximate Bayesian computation" (AABC), a class of methods that extends simulation-based inference by ABC to models in which simulating data is expensive. In AABC, we first simulate a limited number of data sets that is computationally feasible to simulate from the parametric model. We use these data sets as fixed background information to inform a non-mechanistic statistical model that approximates the correct parametric model and enables efficient simulation of a large number of data sets by Bayesian resampling methods. We show that under mild assumptions, the posterior distribution obtained by AABC converges to the posterior distribution obtained by ABC, as the number of data sets simulated from the parametric model and the sample size of the observed data set increase simultaneously. We illustrate the performance of AABC on a population-genetic model of natural selection, as well as on a model of the admixture history of hybrid populations.
Motivation & Objective
- To address the challenge of performing ABC inference when simulating data from a mechanistic model is computationally infeasible due to high cost.
- To develop a method that enables efficient posterior sampling for model-specific parameters in such computationally expensive models.
- To bridge mechanistic modeling with nonparametric Bayesian methods through a two-stage approximation framework.
- To ensure theoretical convergence of the AABC posterior to the true ABC posterior under mild regularity conditions.
Proposed method
- AABC first simulates a limited number of data sets from the computationally expensive mechanistic model, forming a fixed dataset collection.
- It constructs a non-mechanistic statistical model using Dirichlet resampling of these simulated data sets to enable efficient simulation of large numbers of synthetic datasets.
- For each parameter value drawn from the prior, the closest matching parameter in the simulated dataset set is used, introducing a uniform kernel smoothing approximation on the parameter space.
- The method uses Bayesian nonparametric resampling (analogous to Rubin’s Bayesian bootstrap) to assign probabilities to data points from the initial simulations, enabling repeated sampling.
- Posterior inference is performed using likelihoods derived from the non-mechanistic resampled model, not the original mechanistic model.
- The approach allows researchers to fix computation time in advance by setting the number of initial simulations (m), ensuring a desired posterior sample size.
Experimental results
Research questions
- RQ1Can posterior inference be reliably performed when simulating data from a mechanistic model is too expensive for standard ABC?
- RQ2How can a limited number of simulations from a complex model be leveraged to enable large-scale posterior sampling?
- RQ3What is the theoretical relationship between the AABC posterior and the true ABC posterior as the number of initial simulations and observed data size grow?
- RQ4Can non-mechanistic, resampled models preserve the interpretability of model-specific parameters from mechanistic models?
Key findings
- AABC enables posterior sampling for model-specific parameters in models where standard ABC fails due to computational cost of simulating data.
- The posterior distribution obtained via AABC converges to the true ABC posterior as both the number of initial simulations (m) and the observed data size increase.
- For moderate values of m (e.g., 10^3 to 10^4), AABC produces adequate posterior samples where standard ABC methods fail.
- The method provides a practical solution for researchers to pre-specify computation time by fixing m, ensuring a target posterior sample size.
- AABC maintains the interpretability of model-specific parameters from mechanistic models, unlike purely nonparametric approaches.
- The method introduces two approximations: one on the parameter space via nearest-neighbor mapping, and one on the model space via Dirichlet resampling of simulated data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.