Skip to main content
QUICK REVIEW

[Paper Review] Negative binomial models for development triangles of counts

Luis E. Nieto‐Barajas, Rodrigo S. Targino|arXiv (Cornell University)|Jan 9, 2026
Probability and Risk Models0 citations
TL;DR

The paper proposes negative binomial models for run-off triangles of claim counts, introducing dependence across development years via a moving-average latent structure, and provides Bayesian inference with simulation studies and real-data applications.

ABSTRACT

Prediction of outstanding claims has been done via nonparametric models (chain ladder), semiparametric models (overdispersed poisson) or fully parametric models. In this paper, we propose models based on negative binomial distributions for the prediction of outstanding number of claims, which are particularly useful to account for overdispersion. We first assume independence of random variables and introduce appropriate notation. Later, we generalise the model to account for dependence across development years. In both cases, the marginal distributions are negative binomials. We study the properties of the models and carry out bayesian inference. We illustrate the performance of the models with simulated and real datasets.

Motivation & Objective

  • Motivate claim reserving for IBNR counts and extend beyond independent Poisson/ NB setups.
  • Propose a NB development-triangle model with moving-average dependence across development years.
  • Develop a Bayesian inference framework with data augmentation and MCMC.
  • Evaluate model performance on simulated and real insurance datasets.
  • Compare dependence models to independent benchmarks and chain-ladder predictions.

Proposed method

  • Use NB(α_i, 1/(1+π_j)) marginals with row totals α_i and development-year proportions π_j.
  • Introduce a dependence sequence via latent Z and Y with a moving-average structure of order q; X_{i,j} marginally NB(α_i, 1/(1+π_j)).
  • Derive conditional distributions and autocovariances; show Corr(X_{i,j}, X_{i,j+k}) as a function of γ_i,j and π_j.
  • Establish a Bayesian framework with priors α_i ~ Geo(p_α), γ_j ~ Ga(a_γ,b_γ), π ~ Dir(a); augment likelihood with latent Z,Y and use Gibbs sampling with Metropolis-Hastings steps.
  • Assess model fit via LPML, BIAS, and PVAR; implement MCMC with random-walk proposals and tuning for 30% acceptance.
  • Apply to simulated data and real datasets (general insurance and automobile) to select order of dependence q and to compare with chain-ladder.
Figure 2 : Simulated data. Posterior estimates of parameters: $\alpha_{i}$ , $i=1,\ldots,n$ (left) and $\pi_{j}$ , $j=1,\ldots,n$ (right) with $n=10$ . True value (dots) and 95% CI (lines).
Figure 2 : Simulated data. Posterior estimates of parameters: $\alpha_{i}$ , $i=1,\ldots,n$ (left) and $\pi_{j}$ , $j=1,\ldots,n$ (right) with $n=10$ . True value (dots) and 95% CI (lines).

Experimental results

Research questions

  • RQ1Can a negative binomial run-off-triangle model with dependence across development years capture overdispersion and within-triangle correlations?
  • RQ2How does the dependence order q and the strength parameters γ influence model fit and reserve predictions?
  • RQ3Does incorporating development-year dependence improve predictive accuracy and reduce reserve overestimation compared to independence or chain-ladder?
  • RQ4What is the impact of the model on posterior estimates of α_i, π_j, and γ_j across different datasets?

Key findings

  • The dependence model remains NB marginally but introduces cross-year dependence through q and γ_j.
  • Autocorrelation between development years is positive and controlled by γ and π, increasing with stronger dependence and smaller lag.
  • In simulations, LPML, BIAS, and PVAR correctly identify the true order q (e.g., q=2 in the study).
  • On general insurance data, LPML and PVAR favor a dependent model (q=1) while BIAS may favor independence (q=0); nonetheless, dependent models alter posterior summaries.
  • Predictions under the dependent model align with observed development patterns and can produce narrower or shifted reserve estimates than chain-ladder, avoiding overestimation in some cases.
  • Across automobile data, the best-fit model prefers q=1, with posterior predictions for individual N_i and total N providing plausible intervals and sometimes tighter than chain-ladder.
Figure 3 : Simulated data. Left: posterior estimates for parameters $\gamma_{j}$ , $j=1,\ldots,n$ with $n=10$ . True value (dots) and 95% CI (lines). Right: boxplots of posterior predicted aggregated number of claims $N_{i}$ , for $i=2,\ldots,n$ .
Figure 3 : Simulated data. Left: posterior estimates for parameters $\gamma_{j}$ , $j=1,\ldots,n$ with $n=10$ . True value (dots) and 95% CI (lines). Right: boxplots of posterior predicted aggregated number of claims $N_{i}$ , for $i=2,\ldots,n$ .

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.