Skip to main content
QUICK REVIEW

[Paper Review] On Bayesian quantile regression and outliers

Bruno Santos, Heleno Bolfarine|arXiv (Cornell University)|Jan 27, 2016
Bayesian Methods and Mixture Models10 references3 citations
TL;DR

This paper proposes a Bayesian quantile regression approach that uses the posterior distribution of latent variables from an asymmetric Laplace mixture representation to identify outliers in regression models. By treating the scale parameter σ as estimable rather than fixed, the method detects observations deviating significantly from the conditional distribution, especially in different quantiles, as demonstrated in simulations and Brazilian Gini index data with multiple outliers in distinct quantile regions.

ABSTRACT

In this work we discuss the progress of Bayesian quantile regression models since their first proposal and we discuss the importance of all parameters involved in the inference process. Using a representation of the asymmetric Laplace distribution as a mixture of a normal and an exponential distribution, we discuss the relevance of the presence of a scale parameter to control for the variance in the model. Besides that we consider the posterior distribution of the latent variable present in the mixture representation to showcase outlying observations given the Bayesian quantile regression fits, where we compare the posterior distribution for each latent variable with the others. We illustrate these results with simulation studies and also with data about Gini indexes in Brazilian states from years with census information.

Motivation & Objective

  • To address the limitations of fixed scale parameters (σ) in Bayesian quantile regression, which can distort variance control and inference.
  • To develop a robust method for identifying outliers by analyzing the posterior distribution of latent variables in the asymmetric Laplace mixture representation.
  • To demonstrate that multiple outliers can affect different quantiles differently, requiring quantile-specific outlier detection.
  • To provide a computationally feasible alternative to case-deletion diagnostics by leveraging posterior latent variable distributions.
  • To show that fixing σ leads to unreliable posterior covariance and biased estimates, especially in high-dimensional or contaminated data.

Proposed method

  • Uses a scale-mixture representation of the asymmetric Laplace distribution, expressing it as a normal variance-mean mixture with an exponential mixing variable.
  • Models the latent variable $ v_i $ for each observation $ i $, whose posterior distribution reflects the deviation of the residual from the conditional quantile.
  • Applies MCMC sampling to estimate the joint posterior of regression coefficients, $ \sigma $, and latent variables $ v_i $, enabling full Bayesian inference.
  • Compares the posterior distributions of $ v_i $ across observations to identify those with extreme values, indicating potential outliers.
  • Employs the Kullback-Leibler divergence between the posterior of $ v_i $ and the average posterior to quantify outlyingness.
  • Uses mean posterior probability of being an outlier as a diagnostic measure, with higher values indicating greater deviation.

Experimental results

Research questions

  • RQ1How does fixing the scale parameter σ in Bayesian quantile regression affect posterior inference and outlier detection?
  • RQ2Can the posterior distribution of the latent variable $ v_i $ serve as a reliable indicator of outlying observations in Bayesian quantile regression?
  • RQ3How do multiple outliers affect quantile regression estimates across different quantiles, and can they be detected separately?
  • RQ4What is the impact of estimating σ from the posterior versus fixing it at 1 in terms of model fit and inference accuracy?
  • RQ5Can the proposed method detect outliers that are influential in different parts of the conditional distribution, such as lower and upper quantiles?

Key findings

  • The posterior distribution of the latent variable $ v_i $ effectively identifies observations that are distant from the bulk of the data, particularly in the lower and upper quantiles.
  • Observations from the Federal District in Brazil (2010, 2000, 1991) were identified as outliers in the 0.1th quantile due to high income and education, despite low Gini index expectations.
  • The state of Santa Catarina (2010) was flagged as an outlier in the 0.9th quantile due to its unusually low Gini index (0.49) compared to the next lowest (0.53).
  • The method detected that outliers affect regression estimates differently across quantiles, with increased bias observed in simulations when multiple outliers were present.
  • Fixing σ at 1 leads to unreliable posterior covariance matrices and poor inference, while estimating σ from the posterior improves model robustness and variance control.
  • The Kullback-Leibler divergence and mean posterior probability of being an outlier provided consistent and interpretable measures for identifying extreme observations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.