Skip to main content
QUICK REVIEW

[论文解读] A Computational Model of Early Word Learning from the Infant's Point of View

Satoshi Tsutsui, Arjun Chandrasekaran|arXiv (Cornell University)|Jun 4, 2020
Advanced Image and Video Retrieval Techniques参考文献 26被引用 11
一句话总结

本研究提出了首个计算模型,用于模拟早期词汇学习过程,该模型利用婴儿在自然玩具游戏期间的原始第一人称视角视频与注视数据,训练卷积神经网络(CNN)将物体名称与视觉参照物关联起来。模型表明,来自婴儿视角的视觉输入——尤其是对大型物体的持续关注——能够实现成功的词汇学习,为婴儿在真实环境中克服指代不确定性提供了机制性解释。

ABSTRACT

Human infants have the remarkable ability to learn the associations between object names and visual objects from inherently ambiguous experiences. Researchers in cognitive science and developmental psychology have built formal models that implement in-principle learning algorithms, and then used pre-selected and pre-cleaned datasets to test the abilities of the models to find statistical regularities in the input data. In contrast to previous modeling approaches, the present study used egocentric video and gaze data collected from infant learners during natural toy play with their parents. This allowed us to capture the learning environment from the perspective of the learner's own point of view. We then used a Convolutional Neural Network (CNN) model to process sensory data from the infant's point of view and learn name-object associations from scratch. As the first model that takes raw egocentric video to simulate infant word learning, the present study provides a proof of principle that the problem of early word learning can be solved, using actual visual data perceived by infant learners. Moreover, we conducted simulation experiments to systematically determine how visual, perceptual, and attentional properties of infants' sensory experiences may affect word learning.

研究动机与目标

  • 使用婴儿第一人视角的原始感官输入来建模早期词汇学习,而非预处理或符号化数据。
  • 探究婴儿实时感官体验中的视觉、感知与注意力特性如何影响词汇学习。
  • 提供计算层面的可行性证明,表明通过婴儿视角的实际视觉数据可解决早期词汇学习中的指代不确定性问题。

提出的方法

  • 使用头戴式摄像机与眼动追踪设备,在自然玩具游戏过程中收集婴儿与父母的自我视角视频与注视追踪数据。
  • 提取父母命名物体时刻附近的图像帧,构建包含视觉场景与语音标签配对的数据集。
  • 在原始视觉帧上训练基于ResNet的卷积神经网络(CNN),以学习语音物体名称与视觉物体外观之间的关联。
  • 基于物体大小与注意力分布(持续关注 vs. 分散关注)等视觉特性对命名物体实例进行分类。
  • 通过在每次命名事件中从干扰项中正确识别目标物体的准确率来评估模型性能。
  • 开展模拟实验,以分离物体大小与注意力焦点对学习结果的影响。

实验结果

研究问题

  • RQ1深度学习模型是否能仅使用婴儿真实世界互动中的原始第一人称视角视频与注视数据,成功学习物体名称关联?
  • RQ2目标物体在视野中的大小如何影响模型学习词-物映射的能力?
  • RQ3在命名事件中对目标物体的持续关注是否相比分散关注带来更好的学习表现?
  • RQ4婴儿感官输入的视觉与注意力特性在多大程度上有助于解决早期词汇学习中的指代不确定性?

主要发现

  • 当婴儿表现出持续关注时,模型在大型目标物体命名事件中的平均准确率达到24.27%,显著高于小型目标物体的12.18%(p < 0.001)。
  • 在注意力分散的事件中,模型对大型目标的准确率为17.07%,小型目标为12.88%(p < 0.05),证实物体大小对学习结果具有显著影响。
  • 与分散注意力相比,命名事件中持续关注显著提升了学习准确率,支持其在增强视觉输入质量方面的作用。
  • 结果表明,婴儿可获得的视觉输入——尤其是物体大小与注意力焦点——对词汇学习的成功具有直接且可测量的影响。
  • 本研究首次提供了计算证据,表明通过婴儿视角的真实世界感官数据可解决指代不确定性问题。
  • 研究结果支持一种感官层面的解释,即持续关注与显著视觉特征(如大型物体)是早期词汇发展强预测因子的原因。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。