[论文解读] Information Theory and the Length Distribution of all Discrete Systems
本文提出,哈特利-夏农信息守恒(CoHSI)可解释在多种离散系统(如蛋白质、软件功能和音乐作品)中观察到的普遍幂律长度分布。通过证明具有共享字母表的异质系统会产生单峰且幂律尾部的分布,而具有每个组件唯一标记的同质系统则遵循齐普夫定律,从而解释了这一现象。该理论在多个数据集中得到高度统计显著的验证。
We begin with the extraordinary observation that the length distribution of 80 million proteins in UniProt, the Universal Protein Resource, measured in amino acids, is qualitatively identical to the length distribution of large collections of computer functions measured in programming language tokens, at all scales. That two such disparate discrete systems share important structural properties suggests that yet other apparently unrelated discrete systems might share the same properties, and certainly invites an explanation. We demonstrate that this is inevitable for all discrete systems of components built from tokens or symbols. Departing from existing work by embedding the Conservation of Hartley-Shannon information (CoHSI) in a classical statistical mechanics framework, we identify two kinds of discrete system, heterogeneous and homogeneous. Heterogeneous systems contain components built from a unique alphabet of tokens and yield an implicit CoHSI distribution with a sharp unimodal peak asymptoting to a power-law. Homogeneous systems contain components each built from just one kind of token unique to that component and yield a CoHSI distribution corresponding to Zipf's law. This theory is applied to heterogeneous systems, (proteome, computer software, music); homogeneous systems (language texts, abundance of the elements); and to systems in which both heterogeneous and homogeneous behaviour co-exist (word frequencies and word length frequencies in language texts). In each case, the predictions of the theory are tested and supported to high levels of statistical significance. We also show that in the same heterogeneous system, different but consistent alphabets must be related by a power-law. We demonstrate this on a large body of music by excluding and including note duration in the definition of the unique alphabet of notes.
研究动机与目标
- 解释生物演化系统(如蛋白质)与人类设计系统(如软件功能)之间长度分布的显著相似性,尽管其起源、时间尺度和精度截然不同。
- 建立一个基于哈特利-夏农信息守恒(CoHSI)的统一理论框架,适用于异质和同质离散系统。
- 证明观察到的组件长度分布中的幂律行为并非偶然,而是基本信息论原理的必然结果。
- 使用可复现的开源方法,在包括蛋白质组、软件、音乐、语言和元素丰度在内的多种系统中验证该理论。
- 证明同一系统中不同字母表(例如,包含或不包含音符时值的音乐)之间存在幂律关系。
提出的方法
- 将CoHSI嵌入经典统计力学框架,推导由符号标记构建的离散系统中组件长度的分布。
- 区分异质系统(组件间共享字母表)与同质系统(每个组件使用唯一的标记集合),二者分别产生不同的分布类型。
- 应用CoHSI框架,预测异质系统产生单峰且幂律尾部的分布,而同质系统则产生符合齐普夫定律的分布。
- 对互补累积分布函数(ccdf)进行线性回归,以检验幂律行为,报告决定系数(R-squared)和p值以评估统计显著性。
- 通过开源软件包进行可复现性检查,确保所有结果、表格和图表均可在Linux系统上完整重现。
- 通过包含或排除音符时值来重新定义音乐中的唯一字母表,证明同一系统中不同字母表之间存在幂律关系。
实验结果
研究问题
- RQ1为何蛋白质和计算机功能的长度分布——尽管起源、时间尺度和精度截然不同——却表现出几乎相同的单峰幂律尾部形状?
- RQ2是否存在单一的信息论原理,能够解释生物演化与人类设计的离散系统中相似长度分布的出现?
- RQ3异质系统(共享字母表)与同质系统(每个组件使用唯一标记集合)在长度分布模式上存在何种差异?
- RQ4观察到的组件长度分布中的幂律尾部在多大程度上反映了独立于意义或功能的信息守恒定律?
- RQ5同一系统中不同字母表(如包含与不包含时值的音乐)之间是否如理论预测的那样存在幂律关系?
主要发现
- 8000万个UniProt蛋白质和8000万个C语言功能的长度分布均表现出明显的单峰峰值和幂律尾部,蛋白质字母表大小尾部(21.0–30.0)的ccdf p值为6.576×10⁻¹²,决定系数(R-squared)为0.9951。
- 使用共享字母表的异质系统产生具有尖锐单峰峰值并渐近趋近幂律的CoHSI分布,而同质系统则产生与齐普夫定律匹配的分布。
- 该理论预测,同一系统中不同字母表之间必须存在幂律关系,这一预测在音乐中通过比较包含与不包含音符时值的分布得到证实。
- CoHSI框架以高度统计显著的方式解释了包括蛋白质组、软件、音乐、语言和元素丰度在内的多种系统中的观察分布。
- 可复现性包已发布于 http://leshatton.org/index_RE.html,可完整重建所有结果、表格和图表,包括机器环境检查和回归测试。
- 该理论对标记内容不敏感,即其仅依赖于选择数量(字母表大小),而不依赖于标记的功能或语义意义。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。