Skip to main content
QUICK REVIEW

[Paper Review] A novel, divergence based, regression for compositional data

Michail Tsagris|arXiv (Cornell University)|Nov 24, 2015
Geochemistry and Geologic Mapping10 references3 citations
TL;DR

This paper proposes a novel divergence-based regression method for compositional data using a special case of the Jensen-Shannon divergence, which naturally handles zero values without imputation. The method, called ES-OV regression, outperforms Aitchison's log-ratio regression in datasets with many zeros and is recommended for prediction when zero values are prevalent.

ABSTRACT

In compositional data, an observation is a vector with non-negative components which sum to a constant, typically 1. Data of this type arise in many areas, such as geology, archaeology, biology, economics and political science amongst others. The goal of this paper is to propose a new, divergence based, regression modelling technique for compositional data. To do so, a recently proved metric which is a special case of the Jensen-Shannon divergence is employed. A strong advantage of this new regression technique is that zeros are naturally handled. An example with real data and simulation studies are presented and are both compared with the log-ratio based regression suggested by Aitchison in 1986.

Motivation & Objective

  • To develop a new regression model for compositional data that handles zero values naturally, avoiding the need for zero imputation.
  • To address the limitations of traditional log-ratio-based regression (e.g., Aitchison regression) when applied to compositional data with many zero components.
  • To propose a method based on a recently established metric derived from the Jensen-Shannon divergence, suitable for probability distributions and compositional data.
  • To evaluate the performance of the new method against established approaches using real data and simulation studies.
  • To provide a practical, prediction-oriented regression tool for compositional data with high zero-frequency components.

Proposed method

  • The method employs a divergence-based loss function derived from the Jensen-Shannon divergence, specifically a metric proposed by Endres & Schindelin (2003) and Österreicher & Vajda (2003).
  • The regression minimizes the sum of the Kullback-Leibler divergences between observed and fitted compositions, using the midpoint distribution as a reference.
  • The model uses a log-linear link function to map linear predictors to the simplex, ensuring fitted values remain on the probability simplex.
  • Initial values for optimization are obtained via ordinary least squares on the log-ratio transformed data, followed by iterative minimization using the nlm function in R.
  • The method is implemented in R with a custom function 'esov.compreg' that returns estimated coefficients, deviance, and fitted values.
  • The approach avoids the use of log-ratios, thus eliminating issues arising from zero components in the data.

Experimental results

Research questions

  • RQ1Can a divergence-based regression model be developed that naturally handles zero values in compositional data without requiring imputation?
  • RQ2How does the performance of the proposed ES-OV regression compare to Aitchison’s log-ratio regression in the presence of many zero components?
  • RQ3Does the new method provide more accurate predictions than existing approaches when compositional data contain a high proportion of zeros?
  • RQ4What are the practical implications of using a divergence-based loss function derived from the Jensen-Shannon metric in regression for compositional data?
  • RQ5Is the proposed method suitable for prediction even in the absence of full asymptotic theory?

Key findings

  • The ES-OV regression method handles zero values naturally, avoiding the need for zero imputation techniques required by log-ratio-based methods.
  • In datasets with many zero components, the ES-OV regression outperforms Aitchison regression, particularly when the EM algorithm for zero imputation fails.
  • The method achieves better fit than Aitchison regression in a real data example involving Arctic lake sediment composition with multiple zero values.
  • Simulation studies confirm that the ES-OV method provides more robust estimates than Aitchison regression when zero components are frequent.
  • The method is recommended for prediction purposes, especially when the data contain a high proportion of zeros, despite the lack of full asymptotic theory.
  • The proposed method is competitive with or superior to existing approaches in scenarios where zero values are problematic for traditional log-ratio regression.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.