Skip to main content
QUICK REVIEW

[Paper Review] Bayesian Modeling and MCMC Computation in Linear Logistic Regression for Presence-only Data

Fabio Divino, Natalia Golini|arXiv (Cornell University)|May 6, 2013
Statistical Methods and Bayesian Inference25 references3 citations
TL;DR

This paper proposes a Bayesian hierarchical model with data augmentation for linear logistic regression on presence-only data, using a two-level framework to handle censoring and sampling bias. It introduces a novel MCMC algorithm that estimates regression coefficients and population prevalence without requiring prior knowledge of prevalence, demonstrating robust performance across 24,000 simulated datasets with consistent parameter recovery and accurate prevalence estimation.

ABSTRACT

Presence-only data are referred to situations in which, given a censoring mechanism, a binary response can be observed only with respect to on outcome, usually called extit{presence}. In this work we present a Bayesian approach to the problem of presence-only data based on a two levels scheme. A probability law and a case-control design are combined to handle the double source of uncertainty: one due to the censoring and one due to the sampling. We propose a new formalization for the logistic model with presence-only data that allows further insight into inferential issues related to the model. We concentrate on the case of the linear logistic regression and, in order to make inference on the parameters of interest, we present a Markov Chain Monte Carlo algorithm with data augmentation that does not require the a priori knowledge of the population prevalence. A simulation study concerning 24,000 simulated datasets related to different scenarios is presented comparing our proposal to optimal benchmarks.

Motivation & Objective

  • To address inferential challenges in presence-only data where only presence and covariates are observed, with no direct information on absence.
  • To develop a Bayesian framework that accounts for both censoring and sampling bias through a two-level hierarchical model.
  • To enable MCMC-based inference on regression parameters and population prevalence without assuming known population prevalence.
  • To provide a rigorous formalization of logistic regression under presence-only sampling, clarifying underlying assumptions and identifiability issues.
  • To evaluate performance via a large-scale simulation study comparing the proposed method to benchmark models across diverse scenarios.

Proposed method

  • The method employs a two-level model: a latent binary response Y (presence/absence) and an observed case-control indicator C (presence sampling)
  • It models the joint distribution of Y and C using conditional independence given X, allowing derivation of the likelihood under presence-only sampling.
  • A data augmentation scheme is used to impute missing absence data, enabling full Bayesian inference via MCMC.
  • The MCMC algorithm jointly samples latent Y, regression coefficients β, and prevalence π, using Gibbs sampling with conjugate priors.
  • The model does not require prior knowledge of population prevalence π, which is estimated as part of the posterior inference.
  • The approach uses a logistic link function: logit(Pr(Y=1|x)) = β₀ + β₁x, with a hierarchical prior on β and a beta prior on π.

Experimental results

Research questions

  • RQ1How can a Bayesian hierarchical model be formally constructed to handle presence-only data with unknown population prevalence?
  • RQ2What is the impact of sampling bias and censoring on the identifiability and estimation of logistic regression parameters in presence-only data?
  • RQ3Can a data-augmentation-based MCMC algorithm estimate both regression coefficients and population prevalence without requiring prior knowledge of prevalence?
  • RQ4How does the proposed method compare to optimal benchmarks in terms of bias, coverage, and precision across varying sample sizes and prevalence levels?
  • RQ5What is the role of the case-control design in enabling consistent estimation under presence-only sampling?

Key findings

  • The proposed method accurately recovers true regression coefficients β₀ and β₁ across all sample sizes, with posterior medians converging to the true values as sample size increases.
  • Population prevalence π is consistently estimated with low bias and narrow credible intervals, even when the true prevalence is unknown a priori.
  • For sample sizes ≥500, the posterior medians of β₀ and β₁ were within 0.05–0.08 of the true values, with coverage probabilities near 95%.
  • The MCMC algorithm successfully estimates prevalence without requiring it as a known input, outperforming models that assume fixed or misspecified prevalence.
  • In both scenarios (i) and (ii), the method showed stable convergence and low mean squared error, with parameter estimates approaching the true values as sample size increased.
  • The simulation study of 24,000 datasets confirmed that the method maintains good frequentist properties, including correct coverage and low bias, across diverse data-generating mechanisms.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.