Skip to main content
QUICK REVIEW

[Paper Review] Constructing Likelihood Functions for Interval-valued Random Variables

Xin Zhang, Boris Béranger|arXiv (Cornell University)|Jul 30, 2016
Neural Networks and Applications28 references3 citations
TL;DR

This paper develops a likelihood-based inferential framework for interval-valued data by modeling the underlying real-valued data generating process, enabling direct fitting of parametric models from interval summaries. It overcomes the limitations of uniformity assumptions in existing methods by deriving exact likelihoods through order statistics and aggregation mechanisms, demonstrating improved inference accuracy in simulations and real data.

ABSTRACT

There is a growing need for the ability to analyse interval-valued data. However, existing descriptive frameworks to achieve this ignore the process by which interval-valued data are typically constructed; namely by the aggregation of real-valued data generated from some underlying process. In this article we develop the foundations of likelihood based statistical inference for random intervals that directly incorporates the underlying generative procedure into the analysis. That is, it permits the direct fitting of models for the underlying real-valued data given only the random interval-valued summaries. This generative approach overcomes several problems associated with existing methods, including the rarely satisfied assumption of within-interval uniformity. The new methods are illustrated by simulated and real data analyses.

Motivation & Objective

  • To address the lack of robust statistical inference methods for interval-valued data that account for the true data-generating process.
  • To overcome the common but often unjustified assumption of within-interval uniformity in symbolic data analysis.
  • To develop a likelihood-based framework that enables model fitting directly from interval summaries without requiring individual-level data.
  • To provide a generative approach to interval data analysis that preserves statistical efficiency and accuracy.
  • To extend standard likelihood theory to handle interval-valued observations derived from aggregated real-valued data.

Proposed method

  • Models interval-valued data as summaries of underlying real-valued observations, assuming a known data aggregation mechanism (e.g., min/max, quantiles).
  • Derives the exact likelihood function for interval-valued observations by integrating over the joint distribution of the underlying data points.
  • Uses order statistics and multivariate integration to compute the probability that m real-valued observations fall within specified interval bounds.
  • Expresses the likelihood as a sum over combinatorial configurations of boundary and internal points, accounting for all possible ways the interval bounds can be formed.
  • Derives closed-form expressions for the likelihood under different aggregation functions (e.g., min/max, first and third quartiles) using joint density and cumulative distribution functions.
  • Applies the framework to parametric models such as normal and uniform distributions, enabling maximum likelihood estimation from interval summaries.

Experimental results

Research questions

  • RQ1How can likelihood-based inference be constructed for interval-valued data when only interval summaries are available, without assuming within-interval uniformity?
  • RQ2What is the exact likelihood function for interval-valued observations derived from an underlying real-valued data-generating process?
  • RQ3How does the proposed method compare to standard symbolic data analysis approaches that assume uniformity within intervals?
  • RQ4Can the method accurately recover underlying model parameters (e.g., mean, scale) from interval summaries in the presence of aggregation and measurement uncertainty?
  • RQ5What are the implications of using different aggregation functions (e.g., min/max vs. quartiles) on the resulting likelihood and inference quality?

Key findings

  • The proposed likelihood framework successfully recovers true model parameters (e.g., μ=0, σ=2) from interval summaries with minimal bias, even under complex aggregation mechanisms.
  • Simulations show that fitting a normal model to interval data generated from a uniform distribution yields accurate estimates when the generative likelihood is used, outperforming standard uniformity-based methods.
  • The method achieves good goodness-of-fit, as evidenced by quantile-quantile plots closely aligning with the identity line when the correct aggregation function is used.
  • The inclusion of boundary point configurations in the likelihood derivation significantly improves estimation accuracy compared to approximations ignoring such configurations.
  • The framework is robust to 5% outliers in the data, maintaining stable parameter estimates under contamination.
  • The likelihood expression for m observations in a bivariate interval involves combinatorial terms accounting for all possible ways the interval bounds can be formed by the underlying data points.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.