[Paper Review] Automatic generation of ground truth for the evaluation of obstacle detection and tracking techniques
This paper presents a novel, annotation-free method for automatically generating high-precision ground truth kinematics data for obstacle detection and tracking in autonomous driving. By leveraging RTK-GNSS and IMU-equipped ego- and target vehicles, the approach computes relative positions, velocities, and yaw angles in the ego-vehicle frame, achieving root mean square errors of ≤0.12 m for position and ≤0.30 m/s for velocity.
As automated vehicles are getting closer to becoming a reality, it will become mandatory to be able to characterise the performance of their obstacle detection systems. This validation process requires large amounts of ground-truth data, which is currently generated by manually annotation. In this paper, we propose a novel methodology to generate ground-truth kinematics datasets for specific objects in real-world scenes. Our procedure requires no annotation whatsoever, human intervention being limited to sensors calibration. We present the recording platform which was exploited to acquire the reference data and a detailed and thorough analytical study of the propagation of errors in our procedure. This allows us to provide detailed precision metrics for each and every data item in our datasets. Finally some visualisations of the acquired data are given.
Motivation & Objective
- To eliminate manual annotation in ground truth generation for obstacle detection and tracking evaluation.
- To develop a scalable, automated method for generating precise kinematics data using real-world vehicle platforms.
- To quantify and bound the error propagation in the ground truth generation process to ensure reliability.
- To support the evaluation and benchmarking of perception algorithms such as LiDAR clustering and sensor fusion.
- To enable learning of accurate dynamic motion models from real-world data for improved tracking performance.
Proposed method
- Utilizes two or more vehicles: an ego-vehicle with perception sensors (LiDAR, cameras, RADAR) and a target vehicle with high-precision positioning only.
- Employs RTK-GNSS and high-grade IMU for centimeter-level positioning, fused via a Kalman filter to estimate vehicle states and biases.
- Applies a robust time-synchronization method across distributed vehicles to ensure millisecond-level clock alignment.
- Computes relative kinematics (position, velocity, yaw) of target vehicles in the ego-vehicle’s frame using relative motion equations.
- Derives analytical upper bounds for position and velocity covariance matrices using error propagation models.
- Applies post-processing with accurate ephemeris data and smoothing to further improve positioning accuracy.
Experimental results
Research questions
- RQ1Can ground truth for obstacle detection and tracking be generated without manual annotation or human labeling?
- RQ2What is the maximum achievable precision of automatically generated ground truth using multi-vehicle RTK-GNSS and IMU systems?
- RQ3How do errors in positioning and time synchronization propagate to affect the accuracy of generated ground truth data?
- RQ4To what extent can this method support the evaluation and improvement of perception algorithms like LiDAR clustering and sensor fusion?
- RQ5Can the generated ground truth enable learning of more accurate dynamic motion models than standard preset models?
Key findings
- The method achieves a root mean square error of ≤0.12 m for position and ≤0.30 m/s for velocity in the generated ground truth data.
- The analytical error bounds are derived using extreme values of relative distance (50 m), speed (36 m/s), and yaw rate (1 rad/s), ensuring conservative estimates.
- Under 60-second GNSS outage, the positioning error increases to 0.10 m in horizontal position and 0.07 m in altitude, demonstrating robustness.
- The method enables automatic generation of ground truth without any manual annotation, relying only on sensor calibration and data fusion.
- The ground truth data supports evaluation of LiDAR clustering, sensor fusion (LiDAR-Radar-Camera), and dynamic model learning.
- The approach provides a scalable, repeatable, and precise benchmarking framework for autonomous driving perception systems.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.