[Paper Review] Model-aware Quantile Regression for Discrete Data
This paper proposes a model-aware quantile regression approach for discrete responses (e.g., Poisson, Binomial) by constructing a continuous, distribution-specific interpolation of the quantile function, enabling proper uncertainty quantification and mitigating quantile crossing. The method integrates the true response distribution into the quantile modeling framework, allowing for Bayesian inference via R-INLA and improving interpretability and robustness in disease mapping applications.
Quantile regression relates the quantile of the response to a linear predictor. For a discrete response distributions, like the Poission, Binomial and the negative Binomial, this approach is not feasible as the quantile function is not bijective. We argue to use a continuous model-aware interpolation of the quantile function, allowing for proper quantile inference while retaining model interpretation. This approach allows for proper uncertainty quantification and mitigates the issue of quantile crossing. Our reanalysis of hospitalisation data considered in Congdon (2017) shows the advantages of our proposal as well as introducing a novel method to exploit quantile regression in the context of disease mapping.
Motivation & Objective
- To address the limitations of standard quantile regression when applied to discrete response variables such as counts.
- To overcome the non-bijectivity and non-differentiability issues of quantile functions in discrete distributions like Poisson and Binomial.
- To develop a continuous, model-aware interpolation of the quantile function that preserves interpretability and distributional assumptions.
- To enable proper uncertainty quantification and avoid quantile crossing in Bayesian quantile regression for discrete data.
- To extend the Generalized Linear Mixed Model framework to quantile regression by redefining the link function using the generating distribution.
Proposed method
- Proposes a continuous interpolation of the quantile function based on the true underlying discrete distribution (e.g., Poisson, Binomial), ensuring model-awareness.
- Uses the inverse cumulative distribution function (CDF) of the discrete distribution to define a continuous quantile proxy, avoiding arbitrary jittering.
- Reformulates quantile regression as a generalized linear model by linking the quantile to a linear predictor through a distribution-specific link function.
- Employs the penalized complexity (PC) prior for precision and mixing parameters in the conditional autoregressive (CAR) model for spatial effects.
- Utilizes R-INLA for efficient Bayesian inference, leveraging the full posterior distribution for uncertainty quantification.
- Applies the method to hospitalization data, modeling quantile-specific relative risks and identifying high-risk areas via exceedance probabilities.
Experimental results
Research questions
- RQ1Can a continuous, model-aware interpolation of the quantile function be constructed for discrete distributions to enable valid quantile regression?
- RQ2How can proper uncertainty quantification be achieved in quantile regression for discrete data without relying on jittering or working likelihoods?
- RQ3To what extent does the proposed method mitigate quantile crossing in discrete response settings?
- RQ4Can the proposed framework be integrated into the GLMM framework to unify quantile regression with existing likelihood-based models?
- RQ5How does the model-aware approach improve the detection of high-risk areas in disease mapping compared to point-estimate-based methods?
Key findings
- The model-aware interpolation successfully enables quantile regression on discrete data by constructing a continuous, distribution-specific quantile proxy that avoids the pitfalls of jittering.
- The method provides proper uncertainty quantification through full posterior inference, allowing for exceedance probability calculations to identify high-risk regions.
- The approach mitigates quantile crossing by embedding the true data-generating distribution into the model structure.
- In the reanalysis of hospitalization data, deprivation was found to have a consistently negative impact on self-harm rates, with a posterior mean coefficient of 1.981 at the median quantile.
- Rural status was associated with a lower risk of hospitalization, with a posterior mean coefficient of -0.814, indicating a protective effect.
- The use of exceedance probabilities (e.g., Pr(θ_i^0.25 > 1) > 0.95) for the first quartile relative risk improved the detection of high-risk areas compared to point-estimate-based methods.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.