[论文解读] How Evolution Learns to Generalise: Principles of under-fitting, over-fitting and induction in the evolution of developmental organisation
本文提出,自然选择可通过避免对过去环境的过拟合,使发育组织进化出在新环境中具有泛化能力的特性,这与机器学习中的泛化机制类似。通过基于选择优势和发育成本的调控网络模型,研究证明正则化技术(如L1和L2惩罚)可促进进化出能以少量突变适应未知环境的泛化表型分布。
One of the most intriguing questions in evolution is how organisms exhibit suitable phenotypic variation to rapidly adapt in novel selective environments which is crucial for evolvability. Recent work showed that when selective environments vary in a systematic manner, it is possible that development can constrain the phenotypic space in regions that are evolutionarily more advantageous. Yet, the underlying mechanism that enables the spontaneous emergence of such adaptive developmental constraints is poorly understood. How can natural selection, given its myopic and conservative nature, favour developmental organisations that facilitate adaptive evolution in future previously unseen environments? Such capacity suggests a form of extit{foresight} facilitated by the ability of evolution to accumulate and exploit information not only about the particular phenotypes selected in the past, but regularities in the environment that are also relevant to future environments. Here we argue that the ability of evolution to discover such regularities is analogous to the ability of learning systems to generalise from past experience. Conversely, the canalisation of evolved developmental processes to past selective environments and failure of natural selection to enhance evolvability in future selective environments is directly analogous to the problem of over-fitting and failure to generalise in machine learning. We show that this analogy arises from an underlying mechanistic equivalence by showing that conditions corresponding to those that alleviate over-fitting in machine learning enhance the evolution of generalised developmental organisations under natural selection. This equivalence provides access to a well-developed theoretical framework that enables us to characterise the conditions where natural selection will find general rather than particular solutions to environmental conditions.
研究动机与目标
- 理解自然选择如何在具有短视性和保守性的前提下,偏好那些能增强未来未知环境中可演化性的发育架构。
- 研究发育系统在何种条件下会演化出能在多种选择环境中泛化的表型变异。
- 识别防止对过去环境过拟合的机制,转而促进发现广泛适用的表型解决方案。
- 建立机器学习中过拟合与进化中适应不良的发育固定化之间的机制等价性,以正则化作为统一原理。
提出的方法
- 建模一个包含N个基因的调控网络,其中胚胎表型作为[-1,1]^N中的随机向量生成,并通过固定的发育过程转化为成体表型。
- 将适应度定义为 f_S(P_a) = b - λc,其中b是基于成体表型与目标表型S的内积的选择优势,c是基于调控权重的代价项。
- 采用两种代价函数:调控系数的L1范数(|B|_1)和L2范数平方(|B|_2^2),以惩罚强连接或大量连接,类比于机器学习中的正则化。
- 通过在5000个随机采样的胚胎表型上应用分类与计数(CC)方法,经验估计表型分布,并按与目标模式的接近程度进行分类。
- 使用低差异Sobol序列以减少采样偏差,确保基因型空间的均匀覆盖。
- 通过演化后表型分布与过去或所有潜在目标表型的期望分布之间的卡方误差,量化泛化性能。
实验结果
研究问题
- RQ1在何种条件下,自然选择会演化出能泛化到此前未见过的新环境的发育组织?
- RQ2对调控网络权重施加L1/L2正则化如何影响泛化表型分布的演化?
- RQ3对过去环境的发育固定化在多大程度上阻碍了未来的可演化性?这一问题如何缓解?
- RQ4能否通过模拟形式化并验证机器学习中过拟合与进化中适应不良的发育约束之间的类比?
- RQ5选择环境中结构化的规律在多大程度上促进了适应性发育约束的自发出现?
主要发现
- 通过对调控权重施加L1和L2惩罚的正则化,显著增强了能适应新环境的泛化发育组织的演化。
- 在正则化条件下,演化后表型分布与目标分布之间的卡方误差减小,表明泛化性能得到改善。
- 采用L1或L2正则化的发育系统演化出的表型分布对选择环境变化更具鲁棒性,且达到新最优表型所需的突变更少。
- 本研究证明了机器学习中过拟合与进化中适应不良的发育固定化之间存在机制等价性,正则化作为统一解决方案。
- 在正则化下演化出的系统对过去环境的过度特化程度降低,表明进化可能在无前瞻能力的情况下“学习”泛化。
- 使用Sobol序列进行采样提高了经验分布估计的稳定性和准确性,减少了结果中的随机噪声。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。