Skip to main content
QUICK REVIEW

[论文解读] Presenting an extensive lab- and field-image dataset of crops and weeds for computer vision tasks in agriculture

Michael A. Beck, Chenyi Liu|arXiv (Cornell University)|Aug 12, 2021
Smart Agriculture and AI参考文献 29被引用 9
一句话总结

本论文介绍了两个大规模、公开可用的数据集——120万张实验室培育和54万张田间培育的植物图像,按物种标注,并附带植物年龄、生长阶段和环境条件等元数据。这些数据集支持农业计算机视觉研究,其中1.4万张图像的开放访问子集可实现物种分类和杂草-作物区分的快速模型原型设计。

ABSTRACT

We present two large datasets of labelled plant-images that are suited towards the training of machine learning and computer vision models. The first dataset encompasses as the day of writing over 1.2 million images of indoor-grown crops and weeds common to the Canadian Prairies and many US states. The second dataset consists of over 540,000 images of plants imaged in farmland. All indoor plant images are labelled by species and we provide rich etadata on the level of individual images. This comprehensive database allows to filter the datasets under user-defined specifications such as for example the crop-type or the age of the plant. Furthermore, the indoor dataset contains images of plants taken from a wide variety of angles, including profile shots, top-down shots, and angled perspectives. The images taken from plants in fields are all from a top-down perspective and contain usually multiple plants per image. For these images metadata is also available. In this paper we describe both datasets' characteristics with respect to plant variety, plant age, and number of images. We further introduce an open-access sample of the indoor-dataset that contains 1,000 images of each species covered in our dataset. These, in total 14,000 images, had been selected, such that they form a representative sample with respect to plant age and ndividual plants per species. This sample serves as a quick entry point for new users to the dataset, allowing them to explore the data on a small scale and find the parameters of data most useful for their application without having to deal with hundreds of thousands of individual images.

研究动机与目标

  • 解决数字农业中的关键瓶颈:缺乏大规模、高质量、带标注的植物图像数据,以训练机器学习模型。
  • 提供一份全面、富含元数据的作物与杂草数据集,涵盖加拿大草原地区及美国北部各州常见的植物种类,以支持精准农业研究。
  • 使研究人员能够利用真实世界中多样化的植物外观,训练并评估用于物种分类、杂草检测和植物表型分析的计算机视觉模型。
  • 通过精心筛选的1.4万张图像子集,实现快速原型设计,该子集完整保留了作物与杂草在年龄和物种分布上的特征。
  • 通过遵循数据共享最佳实践,发布数据时采用开放获取原则,附带完整元数据和数据表(datasheet),确保数据的长期可访问性与可重复性。

提出的方法

  • 在受控实验室条件下,使用机器人系统自动采集图像,从多个角度(俯视、侧视、斜视)拍摄植物,光照条件一致。
  • 通过安装在拖拉机上的双目立体相机采集田间图像,以俯视视角记录田间作业过程中的视频,并提取帧图像形成田间数据子集。
  • 使用自动化检测算法在图像中生成单株植物的边界框,从而从全幅图像中裁剪出单株植物图像。
  • 实施分层抽样策略,从实验室数据中每种植物各选取1,000张图像,确保植物年龄和个体多样性在样本中成比例体现。
  • 收集元数据,包括植物物种、年龄(以天为单位)、生长阶段以及环境条件(如温度、湿度、相机高度)等,适用于两个数据集。
  • 通过CyVerse和EMILI数据门户分阶段发布数据,完整数据集和子集均以开放获取许可发布,并附带数据表以确保透明性。

实验结果

研究问题

  • RQ1如何系统性地收集并组织大规模、多样化且标注详尽的植物图像数据集,以支持农业计算机视觉研究?
  • RQ21.4万张图像的代表性子集在多大程度上保留了超过120万张图像的完整数据集中的年龄分布与物种分布特征?
  • RQ3在物种多样性、图像数量和元数据丰富度方面,实验室与田间培育植物图像的关键特征是什么?
  • RQ4开放获取的数据共享与元数据标准化在多大程度上能提升数字农业研究的可重复性与创新性?
  • RQ5对于初次接触植物图像数据集的农业应用研究人员,精心筛选的小规模子集具有哪些实际优势?

主要发现

  • 实验室数据集包含超过120万张图像,涵盖14种作物与杂草,均在室内培育,图像从多个角度拍摄,并附带植物年龄与个体身份的元数据。
  • 田间数据集包含54万张来自真实农田的俯视图像,采集于2019年和2020年生长季,额外包含天气与相机高度等元数据。
  • 1.4万张图像的子集完整保留了完整实验室数据集的年龄分布与个体植物代表性,每种植物均选取1,000张图像,确保覆盖均衡。
  • 该子集可通过DOI https://doi.org/10.25739/rwcw-ex45 在CyVerse数据存储库中获取,支持立即用于模型训练与评估。
  • 完整数据集托管于EMILI数据门户 http://emilicanada.com/(数字农业资产地图),数据持续收集,并正扩展至三维点云与高光谱成像。
  • 数据集随附数据表并遵循数据共享最佳实践,显著提升了数字农业研究中的透明度、可重复性与可用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。