Skip to main content
QUICK REVIEW

[论文解读] Environmental drivers of systematicity and generalization in a situated agent

Felix Hill, Andrew K. Lampinen|arXiv (Cornell University)|Oct 1, 2019
Language and cultural evolution参考文献 39被引用 53
一句话总结

本文表明具体现身的多模态代理可以在零-shot 下泛化成分化的语言-地面化行为,且泛化受训练数据多样性、视角约束和在现实环境中的感知丰富性影响。

ABSTRACT

The question of whether deep neural networks are good at generalising beyond their immediate training experience is of critical importance for learning-based approaches to AI. Here, we consider tests of out-of-sample generalisation that require an agent to respond to never-seen-before instructions by manipulating and positioning objects in a 3D Unity simulated room. We first describe a comparatively generic agent architecture that exhibits strong performance on these tests. We then identify three aspects of the training regime and environment that make a significant difference to its performance: (a) the number of object/word experiences in the training set; (b) the visual invariances afforded by the agent's perspective, or frame of reference; and (c) the variety of visual input inherent in the perceptual aspect of the agent's perception. Our findings indicate that the degree of generalisation that networks exhibit can depend critically on particulars of the environment in which a given task is instantiated. They further suggest that the propensity for neural networks to generalise in systematic ways may increase if, like human children, those networks have access to many frames of richly varying, multi-modal observations as they learn.

研究动机与目标

  • 研究标准神经网络架构是否能够在多模态、情境化设置中实现系统性泛化。
  • 确定环境因素如何影响在 grounded 语言任务中成分性理解(动词/名词)的出现。
  • 识别提升对未见对象和动作的零-shot 泛化的关键训练制度因素。

提出的方法

  • 一个具视觉(像素)和语言输入的多模态代理通过一个3层CNN和一个基于LSTM的语言模块来处理观测。
  • 一个基于LSTM的策略和值网络在一个演员-评论家框架内运行,使用分布式执行者进行训练(IMPALA 风格)。
  • 实验测试在 lifting 和 putting 任务中的零-shot 泛化,给谓词绑定到参数,并在不同环境条件下评估泛化。
  • 在3D Unity 环境与2D网格世界之间进行对比分析,比较自我中心视角与分布视角,以及有无语言监督。
  • 对照条件将 vision-language 分类器与情境代理在颜色-形状泛化任务中的表现进行比较。

实验结果

研究问题

  • RQ1一个标准神经架构是否能够在3D交互环境中将动词和名词绑定到新对象并实现泛化?
  • RQ2哪些环境与感知因素促成或阻碍具身代理的系统性泛化?
  • RQ3增加训练多样性(词汇/对象)、限定参照框架、以及更丰富的时序感知是否提升零-shot 泛化?
  • RQ4语言监督在 grounded 任务中的系统性泛化中有多大贡献?

主要发现

  • 代理在零-shot 的提升与放置任务中对新对象和新词-对象配对表现出泛化能力。
  • 在训练中增加词汇/对象的多样性(否定实验)时,泛化有所提高。
  • 自我中心/有界的视觉视角相对于分布视角对泛化有促进作用。
  • 时间维度、丰富的感知输入(在不同视角下移动)比单帧感知更能提升泛化。
  • 在3D环境中,泛化强于相当的2D网格世界设置,语言对训练表现有适度贡献但并非泛化的必需条件。
  • 静态图像上的 vision-language 分类器在测试泛化方面不如具身感知代理,凸显具身感知对系统性泛化的益处。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。