Skip to main content
QUICK REVIEW

[Paper Review] Public Computer Vision Datasets for Precision Livestock Farming: A Systematic Survey

Anil Bhujel, Yibin Wang|arXiv (Cornell University)|Jun 15, 2024
Smart Agriculture and AIAgricultural and Biological Sciences3 citations
TL;DR

This systematic survey identifies and analyzes 58 public computer vision datasets for precision livestock farming (PLF), revealing that nearly half are for cattle, with individual animal detection and color imaging dominating. The study highlights critical gaps in data diversity, annotation quality, and contextual metadata, urging improved dataset curation to advance AI-driven animal health and welfare monitoring.

ABSTRACT

Technology-driven precision livestock farming (PLF) empowers practitioners to monitor and analyze animal growth and health conditions for improved productivity and welfare. Computer vision (CV) is indispensable in PLF by using cameras and computer algorithms to supplement or supersede manual efforts for livestock data acquisition. Data availability is crucial for developing innovative monitoring and analysis systems through artificial intelligence-based techniques. However, data curation processes are tedious, time-consuming, and resource intensive. This study presents the first systematic survey of publicly available livestock CV datasets (https://github.com/Anil-Bhujel/Public-Computer-Vision-Dataset-A-Systematic-Survey). Among 58 public datasets identified and analyzed, encompassing different species of livestock, almost half of them are for cattle, followed by swine, poultry, and other animals. Individual animal detection and color imaging are the dominant application and imaging modality for livestock. The characteristics and baseline applications of the datasets are discussed, emphasizing the implications for animal welfare advocates. Challenges and opportunities are also discussed to inspire further efforts in developing livestock CV datasets. This study highlights that the limited quantity of high-quality annotated datasets collected from diverse environments, animals, and applications, the absence of contextual metadata, are a real bottleneck in PLF.

Motivation & Objective

  • To identify and catalog all publicly available computer vision datasets relevant to precision livestock farming (PLF).
  • To analyze the characteristics, species distribution, imaging modalities, and annotation types across these datasets.
  • To assess the current state of data availability and its implications for developing AI-based livestock monitoring systems.
  • To identify key challenges in dataset quality, diversity, and metadata completeness that hinder PLF innovation.
  • To provide actionable insights for researchers and practitioners to improve dataset development and foster progress in animal welfare and productivity.

Proposed method

  • Conducted a systematic literature review and web-based search across academic, institutional, and open data repositories to identify public livestock CV datasets.
  • Applied predefined inclusion and exclusion criteria to filter datasets based on relevance to computer vision and livestock farming applications.
  • Categorized datasets by species, imaging modality, annotation type, and application domain (e.g., individual detection, behavior recognition).
  • Evaluated dataset quality through metrics such as sample size, annotation consistency, and environmental diversity.
  • Collected and analyzed contextual metadata, including data collection settings, equipment used, and ethical considerations.
  • Synthesized findings into a comprehensive overview of dataset strengths, limitations, and research implications.

Experimental results

Research questions

  • RQ1What is the current landscape of publicly available computer vision datasets for precision livestock farming?
  • RQ2How are these datasets distributed across different livestock species, imaging modalities, and application types?
  • RQ3What are the primary limitations in dataset quality, annotation consistency, and contextual metadata availability?
  • RQ4How do existing datasets support or constrain the development of AI-driven animal health and welfare monitoring systems?
  • RQ5What opportunities exist for improving dataset curation to accelerate innovation in PLF?

Key findings

  • Of the 58 identified public datasets, 48% are focused on cattle, followed by swine (21%), poultry (17%), and other species (14%).
  • Individual animal detection and color imaging are the most prevalent application and imaging modality, respectively.
  • Only a minority of datasets include contextual metadata such as environmental conditions, housing types, or animal health status.
  • Many datasets suffer from limited diversity in animal breeds, rearing environments, and data collection conditions.
  • High-quality, consistently annotated datasets collected across diverse real-world farming settings remain scarce.
  • The absence of standardized metadata and annotation practices presents a significant bottleneck for training robust, generalizable AI models in PLF.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.