Skip to main content
QUICK REVIEW

[Paper Review] 3D-POP -- An automated annotation approach to facilitate markerless 2D-3D tracking of freely moving birds with marker-based motion capture

Hemal Naik, Alex Hoi Hang Chan|arXiv (Cornell University)|Mar 23, 2023
Animal Vocal Communication and BehaviorBiochemistry, Genetics and Molecular Biology3 citations
TL;DR

This paper introduces 3D-POP, a semi-automated method using marker-based motion capture to generate high-quality 2D and 3D keypoint annotations for freely moving birds. By estimating morphological keypoints (e.g., beak, eyes) relative to tracked body markers, the approach produces a large-scale dataset of 300k annotated frames (4M instances) with ground truth for 2D/3D pose, identity, and trajectory, enabling robust markerless 2D-3D tracking and posture estimation in birds.

ABSTRACT

Recent advances in machine learning and computer vision are revolutionizing the field of animal behavior by enabling researchers to track the poses and locations of freely moving animals without any marker attachment. However, large datasets of annotated images of animals for markerless pose tracking, especially high-resolution images taken from multiple angles with accurate 3D annotations, are still scant. Here, we propose a method that uses a motion capture (mo-cap) system to obtain a large amount of annotated data on animal movement and posture (2D and 3D) in a semi-automatic manner. Our method is novel in that it extracts the 3D positions of morphological keypoints (e.g eyes, beak, tail) in reference to the positions of markers attached to the animals. Using this method, we obtained, and offer here, a new dataset - 3D-POP with approximately 300k annotated frames (4 million instances) in the form of videos having groups of one to ten freely moving birds from 4 different camera views in a 3.6m x 4.2m area. 3D-POP is the first dataset of flocking birds with accurate keypoint annotations in 2D and 3D along with bounding box and individual identities and will facilitate the development of solutions for problems of 2D to 3D markerless pose, trajectory tracking, and identification in birds.

Motivation & Objective

  • To address the scarcity of large-scale, high-resolution, multi-view 2D and 3D annotated datasets for freely moving birds in complex social settings.
  • To develop a scalable, semi-automated method for generating accurate 2D and 3D keypoint annotations without manual labeling of morphological features.
  • To enable the training of deep learning models for markerless 2D-3D pose estimation, trajectory tracking, and individual identification in birds.
  • To overcome the challenge of annotating inaccessible morphological keypoints by using reflective markers on accessible body parts as proxies.
  • To create a publicly available benchmark dataset (3D-POP) that supports research in multi-animal tracking and 3D pose estimation under realistic, naturalistic conditions.

Proposed method

  • The method uses a Vicon motion capture system to track reflective markers attached to accessible body parts (e.g., head, backpack) of homing pigeons.
  • It estimates the 3D positions of morphological keypoints (e.g., eyes, beak, tail) by modeling them as rigid body transformations relative to the tracked markers.
  • The approach leverages the assumption that the bird’s head and body behave as rigid bodies, enabling accurate 3D keypoint prediction from marker data.
  • Video data from four synchronized RGB cameras are used to generate 2D projections of the 3D keypoints, forming a multi-view dataset.
  • An outlier detection algorithm is applied to identify and filter noisy annotations, improving data quality.
  • The resulting dataset, 3D-POP, includes 300k annotated frames with 2D/3D keypoint coordinates, bounding boxes, and individual identities for up to 10 birds in a 3.6m × 4.2m arena.

Experimental results

Research questions

  • RQ1Can a motion capture system be used to semi-automatically generate accurate 2D and 3D keypoint annotations for birds without manual labeling of morphological features?
  • RQ2How accurate is the estimation of morphological keypoints using marker-based rigid body assumptions in freely moving birds?
  • RQ3Can deep learning models trained on 3D-POP generalize to markerless pose estimation in real-world bird tracking scenarios?
  • RQ4To what extent does the dataset support multi-animal 2D-3D tracking and identity tracking in complex, naturalistic flocks?
  • RQ5What is the impact of marker-based proxy annotation on the performance of markerless pose estimation models?

Key findings

  • The method achieved a mean PCK05 of 66% and PCK10 of 94% when comparing automatically generated 3D-POP keypoints to manual annotations, indicating strong alignment with ground truth.
  • Only 2.8% of frames involved wing movements that violated the rigid body assumption, validating the method’s robustness for most of the dataset.
  • Pre-trained YOLOv5s and DeepLabCut models trained on 3D-POP successfully predicted keypoint positions in markerless images, demonstrating generalization to markerless inference.
  • The 3D-POP dataset contains approximately 300,000 annotated frames (4 million keypoint instances) across four camera views, with up to 10 individuals per trial.
  • The dataset supports multi-view, multi-individual 2D-3D tracking with ground truth for identity, pose, and trajectory, enabling benchmarking of markerless tracking systems.
  • The method enables large-scale data curation with minimal human labor, significantly reducing the time and effort required for 3D annotation of animal behavior.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.