Skip to main content
QUICK REVIEW

[Paper Review] Monotone function estimation in the presence of extreme data coarsening: Analysis of preeclampsia and birth weight in urban Uganda

Jennifer E. Starling, Catherine Aiken|arXiv (Cornell University)|Dec 14, 2019
Pregnancy and preeclampsia studies59 references4 citations
TL;DR

This paper proposes psBART, a Bayesian nonparametric model that estimates monotone, smooth relationships between gestational age and birth weight while accounting for extreme data coarsening (83% of birth weights rounded to 100g) in low-resource settings. The method reveals pre-eclampsia as the dominant predictor of low birth weight in urban Uganda, with a stronger negative impact at earlier gestational ages, and provides robust uncertainty quantification via posterior inference.

ABSTRACT

This paper proposes a Bayesian hierarchical model to characterize the relationship between birth weight and maternal pre-eclampsia across gestation at a large maternity hospital in urban Uganda. Key scientific questions we investigate include: 1) how pre-eclampsia compares to other maternal-fetal covariates as a predictor of birth weight; and 2) whether the impact of pre-eclampsia on birthweight varies across gestation. Our model addresses several key statistical challenges: it correctly encodes the prior medical knowledge that birth weight should vary smoothly and monotonically with gestational age, yet it also avoids assumptions about functional form along with assumptions about how birth weight varies with other covariates. Our model also accounts for the fact that a high proportion (83%) of birth weights in our data set are rounded to the nearest 100 grams. Such extreme data coarsening is rare in maternity hospitals in high resource obstetrics settings but common for data sets collected in low and middle-income countries (LMICs); this introduces a substantial extra layer of uncertainty into the problem and is a major reason why we adopt a Bayesian approach. Our proposed non-parametric regression model, which we call Projective Smooth BART (psBART), builds upon the highly successful Bayesian Additive Regression Tree (BART) framework. This model captures complex nonlinear relationships and interactions, induces smoothness and monotonicity in a single target covariate, and provides a full posterior for uncertainty quantification. The results of our analysis show that pre-eclampsia is a dominant predictor of birth weight in this urban Ugandan setting, and therefore an important risk factor for perinatal mortality.

Motivation & Objective

  • To model the complex, nonlinear relationship between gestational age and birth weight in a low-resource obstetric setting with extreme data coarsening.
  • To account for prior medical knowledge that birth weight should increase monotonically with gestational age, without assuming a specific functional form.
  • To compare the predictive influence of maternal pre-eclampsia versus other maternal-fetal covariates on birth weight in a setting where such data are poorly characterized.
  • To provide valid uncertainty quantification in the presence of heaped (rounded) birth weight data, common in LMICs but often ignored in standard models.
  • To enable detection of heterogeneous growth patterns across subgroups defined by pre-eclampsia status and infant sex.

Proposed method

  • The authors develop Projective Smooth BART (psBART), a nonparametric Bayesian additive regression tree model that induces smoothness and monotonicity in the target covariate (gestational age) while allowing flexibility in other covariates.
  • The model extends tsBART by incorporating a latent variable framework to handle data coarsening, modeling observed rounded birth weights as interval-censored observations.
  • It uses a Gaussian process prior on the tree structure to enforce smoothness and monotonicity in gestational age, ensuring biologically plausible growth curves.
  • The method employs a full Bayesian posterior inference approach, enabling uncertainty quantification even under extreme data coarsening.
  • The model is implemented via Markov Chain Monte Carlo (MCMC) sampling, with posterior predictive checks and model diagnostics to validate fit.
  • The approach allows for subgroup analysis via tree-based partitioning, with posterior densities used to assess the significance of splits across covariates.

Experimental results

Research questions

  • RQ1How does maternal pre-eclampsia compare to other maternal-fetal covariates as a predictor of birth weight in urban Uganda?
  • RQ2Does the impact of pre-eclampsia on birth weight vary across gestational age, and if so, how?
  • RQ3How can we model the monotone, smooth relationship between gestational age and birth weight when birth weights are heavily coarsened (rounded to 100g)?
  • RQ4Can we quantify uncertainty in birth weight predictions under extreme data coarsening while preserving monotonicity and smoothness?
  • RQ5What is the relative importance of pre-eclampsia, infant sex, and other maternal-fetal factors in predicting birth weight in this low-resource setting?

Key findings

  • Pre-eclampsia is the dominant predictor of low birth weight in urban Uganda, with pre-eclamptic patients delivering significantly lower birth weight babies at any given gestational age compared to normotensive patients.
  • The negative impact of pre-eclampsia on birth weight is slightly larger at earlier gestational ages, indicating a more pronounced effect during early to mid-gestation.
  • Infant sex is the second most important predictor, with female babies being lighter than males in both pre-eclamptic and normotensive groups.
  • Other maternal-fetal covariates have minimal influence on birth weight compared to pre-eclampsia and sex, suggesting limited heterogeneity beyond these factors.
  • The model successfully accounts for 83% data coarsening (rounding to nearest 100g) and provides valid posterior uncertainty estimates, enabling reliable inference despite data limitations.
  • Posterior density plots of birth weight across tree nodes show clear separation between pre-eclampsia and non-pre-eclampsia groups, with increasing overlap within subgroups, confirming the model’s ability to detect meaningful subpopulations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.