Skip to main content
QUICK REVIEW

[Paper Review] The Practical Scope of the Central Limit Theorem

David Draper, Erdong Guo|arXiv (Cornell University)|Nov 24, 2021
Stochastic processes and statistical mechanics63 references4 citations
TL;DR

This paper investigates the practical sample size requirements for the Central Limit Theorem (CLT) to yield accurate approximations in real-world data science and statistical inference. Using historical context, Edgeworth and Cornish-Fisher expansions, and case studies in sampling and gambling, it shows that convergence can be slow for skewed or heavy-tailed distributions, and that standard thresholds like n=30 may be insufficient—offering refined criteria based on skewness, kurtosis, and desired accuracy levels.

ABSTRACT

The extit{Central Limit Theorem (CLT)} is at the heart of a great deal of applied problem-solving in statistics and data science, but the theorem is silent on an important implementation issue: extit{how much data do you need for the CLT to give accurate answers to practical questions?} Here we examine several approaches to addressing this issue -- along the way reviewing the history of this problem over the last 290 years -- and we illustrate the calculations with case-studies from finite-population sampling and gambling. A variety of surprises emerge.

Motivation & Objective

  • To address the long-standing practical question of how much data is needed for the CLT to provide reliable approximations in applied statistics and data science.
  • To evaluate the limitations of common rules of thumb (e.g., n ≥ 30) in light of distributional characteristics such as skewness and kurtosis.
  • To provide quantitative, distribution-specific criteria for when the normal approximation becomes accurate enough for practical use.
  • To illustrate the implications of poor CLT convergence in real-world contexts, including finite-population sampling and gambling strategies.
  • To highlight the role of higher-order moments and asymptotic expansions in assessing the accuracy of normal approximations.

Proposed method

  • Employs the Edgeworth expansion to refine the normal approximation by incorporating skewness and kurtosis, improving accuracy for finite samples.
  • Uses the Cornish-Fisher expansion to adjust quantiles of the normal distribution based on cumulants, enabling more accurate tail probability estimates.
  • Analyzes the convergence behavior of the CLT through case studies involving binomial sampling and a red-black roulette game, both with discrete, bounded, and skewed distributions.
  • Applies the continuity correction to improve discrete-approximation accuracy, especially in lattice-distributed settings.
  • Evaluates accuracy using absolute error on the probability scale, with thresholds at p = 0.975, 0.995, and 0.9995 (z-scores: ±1.960, ±2.576, ±3.291).
  • Uses numerical computation and symbolic tools (e.g., Mathematica) to evaluate the joint density function J(z) and its equivalence across different representations, validating theoretical derivations.

Experimental results

Research questions

  • RQ1How large must a sample size be for the Central Limit Theorem to yield sufficiently accurate normal approximations in practical data science applications?
  • RQ2To what extent do skewness and excess kurtosis affect the rate of convergence in the CLT, and how can these be quantitatively corrected?
  • RQ3Why do standard rules of thumb like n ≥ 30 often fail in practice, especially for skewed or heavy-tailed distributions?
  • RQ4How do Edgeworth and Cornish-Fisher expansions improve the accuracy of normal approximations compared to the standard CLT?
  • RQ5What are the implications of slow CLT convergence for false discovery rates and replicability in scientific research, particularly under conventional p = 0.05 thresholds?

Key findings

  • The CLT's convergence can be extremely slow for skewed or heavy-tailed distributions, and sample sizes as large as n = 100 or more may be needed for accurate approximations, even when n ≥ 30 is traditionally considered sufficient.
  • For a binomial distribution with p = 0.5, the normal approximation begins to perform well only beyond n ≈ 30, but for p = 0.1 or p = 0.9, convergence is much slower, requiring n > 100 for acceptable accuracy.
  • The use of the continuity correction significantly improves the accuracy of normal approximations for discrete distributions, especially in the tails.
  • The Cornish-Fisher expansion provides a more accurate quantile approximation than the standard normal z-score, particularly for non-normal distributions with high skewness or kurtosis.
  • In a red-black roulette gambling model, the probability of losing 35 MUs is 0.394, and the most likely win is 1 MU (probability 0.372), illustrating that even with a large number of plays, the distribution of net gain remains highly skewed.
  • The paper demonstrates that for high-confidence intervals (e.g., 0.9995), larger sample sizes are required than for 0.95, especially when using absolute error criteria, contradicting the assumption that larger confidence levels are less sensitive to sample size.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.