Skip to main content
QUICK REVIEW

[Paper Review] Understanding heavy tails in a bounded world or, is a truncated heavy tail heavy or not?

Arijit Chakrabarty, Gennady Samorodnitsky|arXiv (Cornell University)|Jan 19, 2010
Complex Systems and Time Series Analysis27 references4 citations
TL;DR

This paper investigates whether truncated power-law (heavy-tailed) distributions retain key statistical properties of their untruncated counterparts, distinguishing between soft and hard truncation regimes. It proposes consistent estimation of the tail index and statistical tests to classify truncation type, finding that soft truncation preserves heavy-tailed behavior while hard truncation erodes it, with empirical validation on network data sets.

ABSTRACT

We address the important question of the extent to which random variables and vectors with truncated power tails retain the characteristic features of random variables and vectors with power tails. We define two truncation regimes, soft truncation regime and hard truncation regime, and show that, in the soft truncation regime, truncated power tails behave, in important respects, as if no truncation took place. On the other hand, in the hard truncation regime much of "heavy tailedness" is lost. We show how to estimate consistently the tail exponent when the tails are truncated, and suggest statistical tests to decide on whether the truncation is soft or hard. Finally, we apply our methods to two recent data sets arising from computer networks.

Motivation & Objective

  • To determine to what extent truncated power-law tails retain the characteristic features of untruncated heavy tails.
  • To classify truncation regimes as soft or hard based on statistical behavior.
  • To develop consistent estimators for the tail exponent α when tails are truncated.
  • To propose statistical tests to distinguish between soft and hard truncation in empirical data.
  • To validate the framework on real computer network data sets with truncated power-law behavior.

Proposed method

  • Introduces a triangular array model where observations are i.i.d. random vectors H_j with regularly varying tails, truncated at level M_n.
  • Defines two truncation regimes: soft (where tail behavior mimics untruncated power laws) and hard (where heavy-tailed features are lost).
  • Proposes a Hill estimator with tuning parameters β and γ to estimate the tail index α consistently under truncation.
  • Develops test statistics Z_n(A; γ) and Z_n(A) to test for hard truncation and stronger versions of it, respectively.
  • Uses p-values from these test statistics to assess the plausibility of hard truncation hypotheses.
  • Applies the framework to two real data sets: file sizes and object sizes from computer networks, using empirical tail estimation and hypothesis testing.

Experimental results

Research questions

  • RQ1To what extent do truncated power-law tails retain the statistical features of untruncated heavy tails?
  • RQ2Can one distinguish between soft and hard truncation regimes in empirical data?
  • RQ3How can the tail index α be consistently estimated when data are truncated?
  • RQ4What statistical tests can reliably detect whether truncation is soft or hard?
  • RQ5What sample size and tuning parameter choices are required for valid inference under truncation?

Key findings

  • In the soft truncation regime, truncated power tails behave as if no truncation occurred, preserving key heavy-tailed features.
  • In the hard truncation regime, most characteristics of heavy-tailed behavior are lost, particularly in extreme value statistics.
  • The Hill estimator provides consistent estimation of the tail exponent α even under truncation, when properly tuned.
  • Empirical testing on four network data pieces showed no rejection of the hard truncation hypothesis (p-values > 0.25), though piece 4 showed lower p-values.
  • For the Object Sizes data set (2.2×10^7 observations), soft truncation could not be rejected (test statistics Z_n(A1) increased with A1/A), while hard truncation was also not rejected (p-values > 0.36).
  • A stronger hypothesis of hard truncation was rejected only at ε = 0.4 (p = 0.08), suggesting limited evidence for hard truncation in the Object Sizes data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.