[论文解读] 3D-POP -- An automated annotation approach to facilitate markerless 2D-3D tracking of freely moving birds with marker-based motion capture
本文提出3D-POP,一种半自动方法,利用基于标记的运动捕捉技术,为自由移动的鸟类生成高质量的2D和3D关键点标注。通过估计相对于追踪身体标记的形态学关键点(如喙、眼睛),该方法生成了一个大规模数据集,包含30万帧标注(400万实例),提供2D/3D姿态、身份和轨迹的真实标注,从而实现鸟类中鲁棒的无标记2D-3D跟踪与姿势估计。
Recent advances in machine learning and computer vision are revolutionizing the field of animal behavior by enabling researchers to track the poses and locations of freely moving animals without any marker attachment. However, large datasets of annotated images of animals for markerless pose tracking, especially high-resolution images taken from multiple angles with accurate 3D annotations, are still scant. Here, we propose a method that uses a motion capture (mo-cap) system to obtain a large amount of annotated data on animal movement and posture (2D and 3D) in a semi-automatic manner. Our method is novel in that it extracts the 3D positions of morphological keypoints (e.g eyes, beak, tail) in reference to the positions of markers attached to the animals. Using this method, we obtained, and offer here, a new dataset - 3D-POP with approximately 300k annotated frames (4 million instances) in the form of videos having groups of one to ten freely moving birds from 4 different camera views in a 3.6m x 4.2m area. 3D-POP is the first dataset of flocking birds with accurate keypoint annotations in 2D and 3D along with bounding box and individual identities and will facilitate the development of solutions for problems of 2D to 3D markerless pose, trajectory tracking, and identification in birds.
研究动机与目标
- 为复杂社交环境中自由移动鸟类的大型、高分辨率、多视角2D和3D标注数据集的稀缺性提供解决方案。
- 开发一种可扩展的半自动方法,无需手动标注形态学特征,即可生成准确的2D和3D关键点标注。
- 支持训练深度学习模型,实现鸟类的无标记2D-3D姿态估计、轨迹跟踪和个体识别。
- 通过在可访问的身体部位使用反光标记作为代理,克服难以标注的形态学关键点的挑战。
- 创建一个公开可用的基准数据集(3D-POP),支持在真实自然条件下进行多只动物跟踪与3D姿态估计的研究。
提出的方法
- 该方法使用Vicon运动捕捉系统追踪安装在信鸽可访问身体部位(如头部、背包)上的反光标记。
- 通过将形态学关键点(如眼睛、喙、尾部)建模为相对于追踪标记的刚体变换,估计其3D位置。
- 该方法利用鸟类头部和身体作为刚体的行为假设,从而能够从标记数据中准确预测3D关键点。
- 使用四台同步的RGB摄像机采集的视频数据,生成3D关键点的2D投影,形成多视角数据集。
- 应用异常值检测算法识别并过滤噪声标注,提升数据质量。
- 最终生成的数据集3D-POP包含30万帧标注,包含2D/3D关键点坐标、边界框和最多10只鸟的个体身份,实验区域为3.6米×4.2米。
实验结果
研究问题
- RQ1能否利用运动捕捉系统在无需手动标注形态学特征的情况下,半自动地生成鸟类准确的2D和3D关键点标注?
- RQ2在自由移动的鸟类中,基于标记的刚体假设对形态学关键点的估计精度如何?
- RQ3在3D-POP上训练的深度学习模型是否能泛化到真实世界鸟类跟踪场景中的无标记姿态估计?
- RQ4该数据集在复杂自然群体中支持多只动物2D-3D跟踪与身份识别的程度如何?
- RQ5基于标记的代理标注对无标记姿态估计模型性能的影响是什么?
主要发现
- 将自动生成的3D-POP关键点与人工标注对比,平均PCK05为66%,PCK10为94%,表明与真实标注高度一致。
- 仅有2.8%的帧涉及翅膀运动,违反了刚体假设,验证了该方法在数据集大部分情况下的鲁棒性。
- 在3D-POP上预训练的YOLOv5s和DeepLabCut模型在无标记图像中成功预测了关键点位置,证明了其对无标记推理的泛化能力。
- 3D-POP数据集包含约30万帧标注(400万关键点实例),覆盖四个摄像机视角,每组实验最多包含10名个体。
- 该数据集支持多视角、多主体的2D-3D跟踪,提供身份、姿态和轨迹的真实标注,可作为无标记跟踪系统性能评估的基准。
- 该方法实现了大规模数据整理,仅需极少人工干预,显著减少了动物行为3D标注所需的时间与精力。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。