[论文解读] Coding-theorem Like Behaviour and Emergence of the Universal Distribution from Resource-bounded Algorithmic Probability
本文研究了在图灵机通用性之下的资源受限计算模型中,算法概率与普遍分布如何涌现。通过模拟有限状态机、上下文无关文法和线性有界自动机,研究发现即使较弱的模型也能产生与普遍分布高度相关的输出分布,解释了高达60%的简洁性/复杂性偏差——为算法复杂性提供了一种实用且可计算的近似方法,其性能优于传统的熵和压缩方法。
Previously referred to as `miraculous' in the scientific literature because of its powerful properties and its wide application as optimal solution to the problem of induction/inference, (approximations to) Algorithmic Probability (AP) and the associated Universal Distribution are (or should be) of the greatest importance in science. Here we investigate the emergence, the rates of emergence and convergence, and the Coding-theorem like behaviour of AP in Turing-subuniversal models of computation. We investigate empirical distributions of computing models in the Chomsky hierarchy. We introduce measures of algorithmic probability and algorithmic complexity based upon resource-bounded computation, in contrast to previously thoroughly investigated distributions produced from the output distribution of Turing machines. This approach allows for numerical approximations to algorithmic (Kolmogorov-Chaitin) complexity-based estimations at each of the levels of a computational hierarchy. We demonstrate that all these estimations are correlated in rank and that they converge both in rank and values as a function of computational power, despite fundamental differences between computational models. In the context of natural processes that operate below the Turing universal level because of finite resources and physical degradation, the investigation of natural biases stemming from algorithmic rules may shed light on the distribution of outcomes. We show that up to 60\% of the simplicity/complexity bias in distributions produced even by the weakest of the computational models can be accounted for by Algorithmic Probability in its approximation to the Universal Distribution.
研究动机与目标
- 研究在图灵完备性之下的子通用计算模型中,普遍分布与算法概率如何涌现。
- 评估在乔姆斯基层级中不同计算模型下,算法复杂性近似值的收敛性与类似编码定理的行为。
- 评估有限且资源受限的模型是否能在缺乏完整计算能力的情况下,以高保真度近似普遍分布。
- 将资源受限的算法概率与传统估计器(如香农熵和无损压缩)进行性能比较。
- 探讨由于物理或资源限制而低于图灵通用性的自然过程所蕴含的启示。
提出的方法
- 在计算能力逐步提升的条件下,模拟有限状态转移器(FST)、上下文无关文法(CFG)和线性有界自动机(LBA)的输出分布。
- 通过各模型生成字符串的经验频率分布,计算算法概率的近似值。
- 应用算法编码定理,将字符串频率与算法复杂性关联,使用概率的倒数作为复杂性估计器。
- 将这些近似值与小型通用图灵机(TM(5,2))的经验分布进行比较,作为普遍分布的基准。
- 使用等级相关性和数值收敛度量,评估不同模型间分布的相似性。
- 与标准复杂性估计器(香农熵和无损压缩(Compress))进行性能评估。
实验结果
研究问题
- RQ1在低于图灵通用性的资源受限模型中,其输出在多大程度上再现了普遍分布的统计特性?
- RQ2在乔姆斯基层级的不同计算模型中,算法复杂性近似值的收敛速度与一致性如何?
- RQ3对算法概率的有限近似是否能解释弱计算模型中观察到的显著简洁性/复杂性偏差?
- RQ4非停机模型与停机模型在字符串覆盖范围和与普遍分布的分布相似性方面有何差异?
- RQ5在估计算法复杂性方面,资源受限的算法概率相较于香农熵和无损压缩的相对性能如何?
主要发现
- FST、CFG和LBA的输出分布与来自小型图灵机的普遍分布表现出强烈的等级相关性,表明其具有涌现的类似编码定理的行为。
- 即使是最弱的模型——有限状态转移器(FST)——也能通过算法概率近似,解释其输出分布中高达60%的简洁性/复杂性偏差。
- 在线性有界自动机(LBA)中,其对普遍分布的近似最为准确,无论在等级相关性还是数值收敛性方面均表现最佳。
- 尽管计算成本较高,CFG生成的分布与普遍分布高度相关,但其在大规模采样中效率低于LBA。
- 资源受限的算法概率近似在估计算法复杂性方面优于香农熵和无损压缩,尤其在高算法随机性的字符串上表现更优。
- 非停机模型(如LBA)的收敛分布更接近LBA而非完整图灵机,表明其在通用性之下存在自然的渐近极限。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。