[论文解读] Public Computer Vision Datasets for Precision Livestock Farming: A Systematic Survey
本系统性综述识别并分析了58个用于精准畜牧养殖(PLF)的公开计算机视觉数据集,发现近一半的数据集聚焦于牛,个体动物检测和彩色成像占主导地位。研究揭示了数据多样性、标注质量和上下文元数据方面的关键缺口,呼吁提升数据集的整理质量,以推动基于人工智能的动物健康与福利监测发展。
Technology-driven precision livestock farming (PLF) empowers practitioners to monitor and analyze animal growth and health conditions for improved productivity and welfare. Computer vision (CV) is indispensable in PLF by using cameras and computer algorithms to supplement or supersede manual efforts for livestock data acquisition. Data availability is crucial for developing innovative monitoring and analysis systems through artificial intelligence-based techniques. However, data curation processes are tedious, time-consuming, and resource intensive. This study presents the first systematic survey of publicly available livestock CV datasets (https://github.com/Anil-Bhujel/Public-Computer-Vision-Dataset-A-Systematic-Survey). Among 58 public datasets identified and analyzed, encompassing different species of livestock, almost half of them are for cattle, followed by swine, poultry, and other animals. Individual animal detection and color imaging are the dominant application and imaging modality for livestock. The characteristics and baseline applications of the datasets are discussed, emphasizing the implications for animal welfare advocates. Challenges and opportunities are also discussed to inspire further efforts in developing livestock CV datasets. This study highlights that the limited quantity of high-quality annotated datasets collected from diverse environments, animals, and applications, the absence of contextual metadata, are a real bottleneck in PLF.
研究动机与目标
- 识别并整理所有公开可用的与精准畜牧养殖(PLF)相关的计算机视觉数据集。
- 分析这些数据集的特征、物种分布、成像模式及标注类型。
- 评估当前数据可用性状况及其对开发基于人工智能的畜牧监测系统的影响。
- 识别数据集质量、多样性及元数据完整性方面的关键挑战,这些挑战阻碍了PLF的创新。
- 为研究人员和从业者提供可操作的见解,以改进数据集开发,推动动物福利与生产效率的进步。
提出的方法
- 通过学术、机构及开放数据存储库开展系统性文献回顾与网络搜索,识别公开的畜牧计算机视觉数据集。
- 应用预设的纳入与排除标准,根据与计算机视觉及畜牧养殖应用的相关性筛选数据集。
- 按物种、成像模式、标注类型及应用领域(如个体检测、行为识别)对数据集进行分类。
- 通过样本量、标注一致性及环境多样性等指标评估数据集质量。
- 收集并分析上下文元数据,包括数据采集环境、使用设备及伦理考量。
- 将研究发现综合为对数据集优势、局限性及研究意义的全面概述。
实验结果
研究问题
- RQ1当前公开可用的用于精准畜牧养殖的计算机视觉数据集的总体格局如何?
- RQ2这些数据集在不同畜牧物种、成像模式及应用类型中的分布情况如何?
- RQ3在数据集质量、标注一致性及上下文元数据可用性方面,主要存在哪些限制?
- RQ4现有数据集在多大程度上支持或制约了基于人工智能的动物健康与福利监测系统的发展?
- RQ5在提升数据集整理质量方面,存在哪些机会可加速PLF领域的创新?
主要发现
- 在识别出的58个公开数据集中,48%聚焦于牛,其次为猪(21%)、家禽(17%)及其他物种(14%)。
- 个体动物检测是应用中最常见的类型,而彩色成像是使用最广泛的成像模式。
- 仅少数数据集包含环境条件、饲养类型或动物健康状况等上下文元数据。
- 许多数据集在动物品种、饲养环境及数据采集条件方面的多样性有限。
- 在多样化的真实农场环境中收集的高质量、一致标注的数据集仍然稀缺。
- 缺乏标准化的元数据和标注实践,已成为制约在PLF中训练稳健、可泛化AI模型的重大瓶颈。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。