Skip to main content
QUICK REVIEW

[Paper Review] Zipf's law and L. Levin's probability distributions

Yuri I. Manin|arXiv (Cornell University)|Jan 3, 2013
Computability, Logic, AI Algorithms14 references4 citations
TL;DR

This paper proposes that Zipf's law—observed as a power-law frequency distribution in language—arises naturally from L. Levin's a priori probability distribution when combined with Kolmogorov complexity. By showing that word ranking correlates with exponential complexity and that Levin’s distribution assigns probabilities inversely proportional to prefix Kolmogorov complexity, the authors derive Zipf’s law with exponent -1 as a consequence of fundamental algorithmic information theory principles.

ABSTRACT

Zipf's law in its basic incarnation is an empirical probability distribution governing the frequency of usage of words in a language. As Terence Tao recently remarked, it still lacks a convincing and satisfactory mathematical explanation. In this paper I suggest that at least in certain situations, Zipf's law can be explained as a special case of the a priori distribution introduced and studied by L. Levin. The Zipf ranking corresponding to diminishing probability appears then as the ordering determined by the growing Kolmogorov complexity. One argument justifying this assertion is the appeal to a recent interpretation by Yu. Manin and M. Marcolli of asymptotic bounds for error--correcting codes in terms of phase transition. In the respective partition function, Kolmogorov complexity of a code plays the role of its energy. This version contains minor corrections and additions.

Motivation & Objective

  • To provide a mathematical explanation for Zipf’s law, which lacks a fully satisfactory theoretical foundation despite its empirical prevalence.
  • To connect Zipf’s law to algorithmic information theory by linking word frequency to Kolmogorov complexity.
  • To show that the inverse frequency ranking in Zipf’s law corresponds to increasing Kolmogorov complexity up to a multiplicative constant.
  • To argue that the a priori probability distribution of L. Levin, when applied to structured, generative systems, naturally yields Zipfian frequency distributions.
  • To reconcile the principle of 'minimization of effort' with algorithmic complexity by reinterpreting effort as the length of the shortest program describing an object.

Proposed method

  • Define a constructive world of computable functions using a formal system of operations (γ, σ, ρ, μ, ι) and structural numberings.
  • Use Kolmogorov complexity $K(w)$ as a measure of algorithmic information content, with $K(w)$ defined via optimal encoding up to $\exp(O(1))$ factors.
  • Apply L. Levin’s a priori probability distribution, which assigns probability $\sim KP(w)^{-1}$, where $KP(w)$ is the exponentiated prefix complexity.
  • Establish that the ranking of objects by decreasing frequency (Zipf’s law) corresponds to increasing $K(w)$, up to $\exp(O(1))$ factors.
  • Leverage the asymptotic equivalence between $K(w)$ and $KP(w)$, with $K(w) \preceq KP(w) \preceq K(w) \cdot \log^{1+\varepsilon} K(w)$, to justify the power-law behavior.
  • Draw analogies to error-correcting codes and phase transitions in coding theory, where Kolmogorov complexity acts as energy in a partition function.

Experimental results

Research questions

  • RQ1Can Zipf’s law be derived from algorithmic complexity rather than heuristic cost models?
  • RQ2What is the role of Kolmogorov complexity in determining the frequency ranking of linguistic items?
  • RQ3How does L. Levin’s a priori probability distribution relate to the observed power-law behavior in word frequencies?
  • RQ4Why does the exponent in Zipf’s law stabilize at -1, and is this explained by complexity-based priors?
  • RQ5To what extent does the distinction between 'observed' and 'generated' systems affect the emergence of Zipfian distributions?

Key findings

  • Zipf’s law with exponent -1 emerges as a consequence of ranking objects by increasing Kolmogorov complexity, up to $\exp(O(1))$ factors.
  • The L. Levin a priori distribution assigns probability $\sim KP(w)^{-1}$, where $KP(w)$ is the exponentiated prefix complexity, and this leads to a power-law frequency distribution.
  • The slight discrepancy between $K(w)$ and $KP(w)$—bounded by $\log^{1+\varepsilon} K(w)$—explains why the full infinite series $\sum K(m)^{-1}$ diverges, but finite approximations still yield Zipfian behavior.
  • The model reinterprets 'effort' in Zipf’s original minimization principle as the length of the shortest program describing an object, i.e., $K(w)$.
  • The framework applies not only to language but also to generative systems such as error-correcting codes, where complexity plays the role of energy in a thermodynamic-like partition function.
  • The approach unifies Zipf’s law with algorithmic information theory, offering a deeper explanation than heuristic cost models or stochastic assumptions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.