[论文解读] Identification of Probabilities
该论文证明,在可计算性约束下,一个类脑系统几乎必然能够从有限的独立同分布样本中识别出真实的概率模型——无论是独立同分布分布还是马尔可夫链,方法是利用大数定律;对于依赖序列,它能够利用柯尔莫哥洛夫复杂度识别出使该序列成为马丁-洛夫典型序列的可计算测度。关键结果是,即使在无计算或数据限制的情况下,这种识别也能在极限下实现。
Within psychology, neuroscience and artificial intelligence, there has been increasing interest in the proposal that the brain builds probabilistic models of sensory and linguistic input: that is, to infer a probabilistic model from a sample. The practical problems of such inference are substantial: the brain has limited data and restricted computational resources. But there is a more fundamental question: is the problem of inferring a probabilistic model from a sample possible even in principle? We explore this question and find some surprisingly positive and general results. First, for a broad class of probability distributions characterised by computability restrictions, we specify a learning algorithm that will almost surely identify a probability distribution in the limit given a finite i.i.d. sample of sufficient but unknown length. This is similarly shown to hold for sequences generated by a broad class of Markov chains, subject to computability assumptions. The technical tool is the strong law of large numbers. Second, for a large class of dependent sequences, we specify an algorithm which identifies in the limit a computable measure for which the sequence is typical, in the sense of Martin-Lof (there may be more than one such measure). The technical tool is the theory of Kolmogorov complexity. We analyse the associated predictions in both cases. We also briefly consider special cases, including language learning, and wider theoretical implications for psychology.
研究动机与目标
- 确定在理论上是否可能从有限样本中识别出概率模型,即使在计算与数据资源无限的假设下。
- 研究在与认知科学和神经科学一致的前提下,从感官或语言输入中学习可计算概率分布与测度的可行性。
- 建立学习算法几乎必然在极限下识别出正确生成模型的条件,仅依赖于数据流。
- 将该分析扩展至独立同分布过程与马尔可夫链,并进一步推广到更一般的依赖序列。
- 探讨这些结果对感知理论、语言学习以及贝叶斯大脑假说的理论影响。
提出的方法
- 利用大数定律证明,从有限的独立同分布样本中,可计算的独立同分布分布可在极限下被识别。
- 应用可计算性假设,确保所考虑的概率分布与马尔可夫链可通过具有随机访问能力的图灵机表示。
- 对于依赖序列,采用马丁-洛夫随机性与典型性概念,识别出使数据序列成为算法随机的可计算测度。
- 利用柯尔莫哥洛夫复杂度定义典型性检验:若序列 x₁…xⱼ 对测度 μ 具有典型性,则量 σ(j) = K(μ) − K(x₁…xⱼ) 保持有界。
- 构建一个学习算法,通过 σ(j) 的下半可计算逼近,在极限下识别出最小索引 k,使得数据序列对 μₖ 具有典型性。
- 利用可计算测度的共可计算枚举与一种新颖的索引技巧(类比于 h 指数),区分有界与无界 σ(j) 序列,替代了先前有缺陷的分离论证。
实验结果
研究问题
- RQ1在计算与数据资源无限的前提下,是否可能从有限数据样本中识别出真实的概率模型,即使在原则上?
- RQ2能否通过仅使用大数定律,使学习算法几乎必然地从有限的独立同分布样本中识别出可计算的独立同分布概率分布?
- RQ3在可计算性约束下,该识别过程能否扩展至马尔可夫链?
- RQ4对于依赖序列,能否识别出一个可计算测度,使得观测到的数据序列在马丁-洛夫意义下具有典型性?
- RQ5在认知科学中,学习概率模型的理论极限是什么?可计算性与算法复杂度如何约束或促进此类学习?
主要发现
- 对于一大类可计算的独立同分布分布与马尔可夫链,学习算法几乎必然能从有限的独立同分布样本中,在极限下识别出真实分布,方法是利用大数定律。
- 在底层分布为可计算的假设下,识别过程保证随着样本量增加而收敛至正确模型。
- 对于依赖序列,若可计算类中存在使数据序列成为马丁-洛夫典型的测度,则该算法可在极限下识别出这样的可计算测度 μ。
- 关键技术洞见是:有界 σ(j) = K(μ) − K(x₁…xⱼ) 意味着典型性,且该有界性可通过可计算测度的共可计算枚举与基于索引的选择规则检测到。
- 该算法在有限时间 n₀ 后输出一个稳定索引 k,使得对所有 n ≥ n₀,有 iₙ = k,从而证明了极限收敛性。
- 该证明修正了早期版本(arXiv:1208.5003)中的缺陷,通过基于 h 指数类比的稳健索引方法,替代了原先错误的有界与无界 σ(j) 序列分离论证。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。