[Paper Review] Shannon Information and Kolmogorov Complexity
This paper provides a comprehensive comparison of Shannon information theory and Kolmogorov complexity, contrasting their foundational concepts—entropy vs. algorithmic complexity, probabilistic vs. algorithmic mutual information, and rate distortion vs. structure function—while demonstrating that expected algorithmic mutual information equals probabilistic mutual information and that universal codes exist for rich classes of sources, including Markov processes.
We compare the elementary theories of Shannon information and Kolmogorov complexity, the extent to which they have a common purpose, and where they are fundamentally different. We discuss and relate the basic notions of both theories: Shannon entropy versus Kolmogorov complexity, the relation of both to universal coding, Shannon mutual information versus Kolmogorov (`algorithmic') mutual information, probabilistic sufficient statistic versus algorithmic sufficient statistic (related to lossy compression in the Shannon theory versus meaningful information in the Kolmogorov theory), and rate distortion theory versus Kolmogorov's structure function. Part of the material has appeared in print before, scattered through various publications, but this is the first comprehensive systematic comparison. The last mentioned relations are new.
Motivation & Objective
- To systematically compare the core concepts of Shannon information theory and Kolmogorov complexity, highlighting their shared goals and fundamental differences.
- To clarify the relationship between probabilistic notions (e.g., entropy, mutual information) and algorithmic notions (e.g., Kolmogorov complexity, algorithmic mutual information).
- To establish that expected algorithmic mutual information equals probabilistic mutual information, bridging the two theories.
- To relate rate distortion theory in Shannon theory to Kolmogorov’s structure function, showing equivalence in expected form.
- To demonstrate the existence of universal codes for rich classes of sources, such as Markov processes, that achieve optimal expected code lengths.
Proposed method
- Define and compare Shannon entropy and Kolmogorov complexity as measures of information, emphasizing that Shannon entropy depends on source distribution while Kolmogorov complexity depends on the object itself.
- Introduce universal codes via a two-part coding scheme: first encoding the model index (e.g., distribution parameter) using a prefix code, then encoding the data using the optimal code for that model.
- Prove that the expected code length of the two-part universal code is within O(log n) of the Shannon-Fano code length for any source in a given class, satisfying the universal code condition.
- Establish that the expected algorithmic mutual information between two objects equals their probabilistic mutual information, using the universal code framework.
- Relate the rate distortion function in Shannon theory to the expected structure function in Kolmogorov complexity, showing they coincide in expectation.
- Use the universal code construction to show that for i.i.d. Bernoulli sources, the average code length converges to the entropy, proving universality for uncountable classes like biased coins.
Experimental results
Research questions
- RQ1How do Shannon entropy and Kolmogorov complexity compare in their treatment of information content in individual objects versus ensemble sources?
- RQ2What is the relationship between probabilistic mutual information and algorithmic mutual information, and under what conditions do they coincide?
- RQ3Can a universal code be constructed that achieves optimal expected code length across a rich class of sources, such as Markov processes?
- RQ4How does rate distortion theory in Shannon theory relate to the structure function in Kolmogorov complexity theory?
- RQ5To what extent can the two-part universal code framework achieve both individual sequence and average-case optimality?
Key findings
- The expected algorithmic mutual information between two objects equals their probabilistic mutual information, establishing a fundamental bridge between the two theories.
- The two-part universal code achieves expected code length within O(log n) of the optimal Shannon-Fano code for any source in a given class, satisfying the universal code condition.
- For i.i.d. Bernoulli sources with bias p, the average code length of the universal code converges to the entropy H(p,1−p), proving universality even for uncountable source classes.
- The expected structure function in Kolmogorov complexity equals the distortion-rate function in rate distortion theory, showing a deep equivalence between lossy compression and algorithmic sufficient statistics.
- A universal code exists for the class of all Markov sources of each order, enabling real-time encoding and decoding with near-optimal performance.
- The Kolmogorov complexity of a set, function, or distribution can be defined via algorithmic sufficient statistics, capturing meaningful information beyond randomness.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.