[论文解读] Evolving Self-taught Neural Networks: The Baldwin Effect and the Emergence of Intelligence
本文提出了一种持续演化的自教自导神经网络——一种能够通过内在动机自主学习而无需外部监督的神经网络。通过在多智能体觅食环境中结合进化与自教自导机制,该系统实现了优于单独使用进化或自教自导的涌现智能觅食策略,展示了人工神经系统中的巴尔德温效应。
The so-called Baldwin Effect generally says how learning, as a form of ontogenetic adaptation, can influence the process of phylogenetic adaptation, or evolution. This idea has also been taken into computation in which evolution and learning are used as computational metaphors, including evolving neural networks. This paper presents a technique called evolving self-taught neural networks - neural networks that can teach themselves without external supervision or reward. The self-taught neural network is intrinsically motivated. Moreover, the self-taught neural network is the product of the interplay between evolution and learning. We simulate a multi-agent system in which neural networks are used to control autonomous agents. These agents have to forage for resources and compete for their own survival. Experimental results show that the interaction between evolution and the ability to teach oneself in self-taught neural networks outperform evolution and self-teaching alone. More specifically, the emergence of an intelligent foraging strategy is also demonstrated through that interaction. Indications for future work on evolving neural networks are also presented.
研究动机与目标
- 开发一种新型神经网络方法,使其能够在无外部监督或奖励的情况下实现自我教学。
- 研究进化与自教自导之间的相互作用如何增强自主智能体的适应性行为。
- 模拟一个多智能体环境,其中智能体需进行觅食与竞争,以检验智能策略的涌现。
- 通过学习与进化适应的协同作用,展示人工神经网络中的巴尔德温效应。
- 探索自教自导神经网络在无监督学习及未来通用人工智能发展中的潜力。
提出的方法
- 使用遗传算法对神经网络进行演化,以优化初始权重和网络结构。
- 智能体通过内在动机引导自教自导,基于从环境交互中衍生出的内部奖励信号更新权重。
- 系统运行于一个具有有限感官输入、且无地图布局或其它智能体先验知识的多智能体觅食世界中。
- 自教自导通过无监督学习实现,智能体通过内部反馈改进行为,而无需外部标签。
- 进化与自教自导共同演化:自教自导的智能体通过在多代中提升适应度,影响进化轨迹。
- 该方法将进化搜索与固定深度、浅层前馈网络架构中的基于梯度的学习相结合。
实验结果
研究问题
- RQ1神经网络能否在无外部监督的情况下,通过自教自导自主发展出智能觅食策略?
- RQ2进化与自教自导之间的相互作用如何增强适应性行为,使其超越单独使用任一方法的效果?
- RQ3当学习与进化在多智能体系统中协同演化时,人工神经网络中是否会出现巴尔德温效应?
- RQ4自教自导中的内在动机是否能导致在竞争环境中涌现出复杂且协调的行为?
- RQ5与随机初始化相比,演化得到的初始权重在多大程度上能提高自教自导的效率?
主要发现
- EVO+Self-taught 方法在实现觅食成功方面显著优于仅使用进化或仅使用自教自导的方法。
- 智能觅食策略在多代中逐渐涌现,即使在没有先验地图知识的情况下,大多数智能体在最终代也学会了抵达食物源。
- 该系统展示了巴尔德温效应:自教自导引导进化朝向更具适应性的解决方案,而这些解决方案在仅靠进化或学习的情况下均无法出现。
- 最终代的智能体表现出协调行为,尽管感官输入有限且无通信机制,仍能有效导航至食物源并展开竞争。
- 结果表明,演化得到的初始权重能够使自教自导比随机初始化更快、更高效地进行,支持了进化塑造神经网络学习能力的假设。
- 本研究提供了证据,表明自教自导神经网络可作为无监督或弱监督学习的基础,对人工通用智能的发展具有重要意义。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。