Skip to main content
QUICK REVIEW

[论文解读] A Survey on Machine Learning from Few Samples

Jiang Lu, Pinghua Gong|arXiv (Cornell University)|Sep 6, 2020
Domain Adaptation and Few-Shot Learning参考文献 367被引用 24
一句话总结

本篇全面综述回顾了2000年代至2019年间的300余篇少样本学习(FSL)论文,将FSL方法分类为生成式与判别式模型,尤其聚焦于基于元学习的方法。该综述提出了一种类别层次化分类法,分析了元学习策略的演进历程,并强调了新兴主题、基准测试与应用场景,为低数据量人工智能系统领域的未来研究提供了系统性基础。

ABSTRACT

Few sample learning (FSL) is significant and challenging in the field of machine learning. The capability of learning and generalizing from very few samples successfully is a noticeable demarcation separating artificial intelligence and human intelligence since humans can readily establish their cognition to novelty from just a single or a handful of examples whereas machine learning algorithms typically entail hundreds or thousands of supervised samples to guarantee generalization ability. Despite the long history dated back to the early 2000s and the widespread attention in recent years with booming deep learning technologies, little surveys or reviews for FSL are available until now. In this context, we extensively review 300+ papers of FSL spanning from the 2000s to 2019 and provide a timely and comprehensive survey for FSL. In this survey, we review the evolution history as well as the current progress on FSL, categorize FSL approaches into the generative model based and discriminative model based kinds in principle, and emphasize particularly on the meta learning based FSL approaches. We also summarize several recently emerging extensional topics of FSL and review the latest advances on these topics. Furthermore, we highlight the important FSL applications covering many research hotspots in computer vision, natural language processing, audio and speech, reinforcement learning and robotic, data analysis, etc. Finally, we conclude the survey with a discussion on promising trends in the hope of providing guidance and insights to follow-up researches.

研究动机与目标

  • 提供2000年代至2019年少样本学习(FSL)研究的及时且全面的综述,以弥补该领域系统性综述的不足。
  • 基于其基本原理,将FSL方法划分为生成式与判别式建模范式。
  • 强调并系统分析基于元学习的FSL方法,包括五个子类:Learn-to-Measure、Learn-to-Finetune、Learn-to-Parameterize、Learn-to-Adjust与Learn-to-Remember。
  • 总结FSL的新兴扩展,如半监督、无监督及跨域FSL,并回顾这些领域的最新进展。
  • 突出FSL在计算机视觉、自然语言处理、语音、机器人学与医疗健康等领域的实际应用,识别关键基准与开放性挑战。

提出的方法

  • 对2000年代初至2019年间的300余篇FSL论文开展系统性文献综述,涵盖从Congealing模型等基础模型到现代元学习框架的各类方法。
  • 提出一种层次化分类法,将FSL方法划分为生成式与判别式模型,并基于泛化能力与学习原理进一步细分。
  • 通过识别五项核心学习目标,分析基于元学习的FSL:Learn-to-Measure、Learn-to-Finetune、Learn-to-Parameterize、Learn-to-Adjust与Learn-to-Remember。
  • 回顾关键FSL基准测试,如miniImageNet、Omniglot,以及CVPR 2020新推出的跨域少样本学习挑战赛。
  • 使用标准化指标评估FSL在各领域的性能,包括5类1-shot与5-shot任务的少样本准确率。
  • 整合认知科学与神经生物学的洞见,将FSL理论框架化为稀疏数据下的函数正则化问题。

实验结果

研究问题

  • RQ12000年至2019年间,FSL方法如何从早期的生成式模型演进为现代的元学习框架?
  • RQ2在建模原理与泛化机制方面,生成式与判别式FSL方法之间的核心差异与关系是什么?
  • RQ3基于元学习的FSL五种子类(Learn-to-Measure、Learn-to-Finetune等)在学习目标与架构设计上如何不同?
  • RQ4在半监督、无监督及跨域FSL等新兴FSL扩展中,关键挑战与最新进展是什么?
  • RQ5医学、机器人学与自然语言处理等实际应用如何体现FSL在低数据场景下的实用可行性与局限性?

主要发现

  • 综述指出,元学习是现代FSL的主导范式,其中Learn-to-Measure与Learn-to-Finetune是最广泛采用的策略。
  • 在miniImageNet等标准基准上的表现显示,元学习模型在5类1-shot任务上的准确率超过80%,显著优于非元学习基线模型。
  • CVPR 2020举办的跨域少样本学习挑战赛表明,基于ImageNet训练的模型可泛化至皮肤镜图像与卫星图像等多样化领域,但平均性能下降10%–20%。
  • 半监督与无监督FSL扩展展现出潜力,近期方法在利用未标记目标数据时,性能相较监督基线最高可提升15%。
  • 尽管已有进展,领域分布偏移与数据噪声下的泛化能力仍是主要挑战,现有模型在异常值或分布偏移下表现出显著性能下降。
  • 理论分析表明,所有FSL方法均隐式地对函数空间进行正则化,提示可基于稀疏监督下的函数正则化构建统一框架。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。