[论文解读] From Entropy to Information: Biased Typewriters and the Origin of Life
本文利用数字生命计算模型(Avida)研究了自复制信息系统的自发涌现,表明由先前自复制体塑造的指令分布偏差可显著提高发现功能性复制体的可能性。关键发现是,信息丰富且非均匀的序列分布可大幅提高自复制进化的概率,暗示早期生命可能并非源于纯粹的随机性,而是源于前生命化学中熵减少、信息累积的过程。
The origin of life can be understood mathematically to be the origin of information that can replicate. The likelihood that entropy spontaneously becomes information can be calculated from first principles, and depends exponentially on the amount of information that is necessary for replication. We do not know what the minimum amount of information for self-replication is because it must depend on the local chemistry, but we can study how this likelihood behaves in different known chemistries, and we can study ways in which this likelihood can be enhanced. Here we present evidence from numerical simulations (using the digital life chemistry "Avida") that using a biased probability distribution for the creation of monomers (the "biased typewriter") can exponentially increase the likelihood of spontaneous emergence of information from entropy. We show that this likelihood may depend on the length of the sequence that the information is embedded in, but in a non-trivial manner: there may be an optimum sequence length that maximizes the likelihood. We conclude that the likelihood of spontaneous emergence of self-replication is much more malleable than previously thought, and that the biased probability distributions of monomers that are the norm in biochemistry may significantly enhance these likelihoods
研究动机与目标
- 探究在数字演化计算模型中,自复制系统是否能从随机序列中自发涌现。
- 确定由先前自复制体塑造的指令分布偏差如何影响发现新自复制体的可能性。
- 检验信息含量(而非序列长度)是否是自复制自发涌现的关键因素。
- 探讨熵减少与非均匀概率分布在家原生信息系统起源中的作用。
提出的方法
- 在Avida中使用均匀(无偏差)的指令分布(每条指令概率为1/26),生成长度为L = 8、15、30和100的随机基因组。
- 将自复制体定义为能够成功分裂并产生可存活、自复制后代(形成菌落)的生物体,采用Avida的默认寿命和复制规则。
- 引入基于公式的指令分布偏差:$ p(i,b) = (1-b)(1/26) + b p_\star(i) $,其中$ p_\star(i) $为此前发现的自复制体中指令i的频率。
- 执行迭代偏差处理:利用前一次搜索中发现的自复制体来定义下一次的偏差(第一轮、第二轮、第三轮偏差),逐步将指令分布引导至功能性序列。
- 开展大规模搜索:在无偏差条件下,L=8、15、30时各搜索$10^9$条序列,L=100时搜索$3\times10^8$条序列;在有偏差条件下,L=15及L=8、30时每种偏差水平各搜索$10^8$条序列。
- 测量不同偏差水平下自复制体的出现频率,以评估指令分布熵如何影响复制体发现概率。
实验结果
研究问题
- RQ1在不同基因组长度下,随机无偏差搜索中自复制序列自发涌现的可能性有多大?
- RQ2将指令分布偏向先前发现的自复制体中的序列,如何影响发现新自复制体的概率?
- RQ3信息含量(通过熵减少衡量)在多大程度上决定了自复制体涌现的成功,而非序列长度?
- RQ4能否通过迭代偏差指令分布,实现自复制体发现率的累积性提升,从而模拟前生命阶段的信息累积过程?
主要发现
- 随着指令分布偏差的增加,特别是高偏差水平(b=1)时,发现自复制体的概率显著提高,表明非均匀分布可增强复制体的涌现。
- 对于L=15,全偏差(b=1)条件下的自复制体发现数量远高于均匀分布条件,表明功能性序列在序列空间中并非随机分布。
- 迭代偏差处理(第一轮、第二轮、第三轮偏差)导致自复制体频率逐步上升,表明存在一种信息累积的自我强化过程。
- 研究发现,决定自复制体发现的关键因素是信息含量,而非序列长度,因为即使在均匀分布下,更长的序列也未带来更高的成功率。
- 结果支持如下假设:类似生命的信息系统并非完全依赖偶然性,而是通过前生命系统中熵减少、信息累积的过程实现的。
- 该模型表明,具有偏向性的打字机(即某些指令更受青睐)可模拟功能性复杂性的出现,为生命起源提供了一个合理的机制。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。