[论文解读] FLSea: Underwater Visual-Inertial and Stereo-Vision Forward-Looking Datasets
FLSea 提供公开的前向观测的水下立体视觉和视觉-惯性数据集,附带地面真实深度图和经标定的传感器数据,以促进在具有挑战性的水下环境中的 SLAM、VO 和深度估计研究。
Visibility underwater is challenging, and degrades as the distance between the subject and camera increases, making vision tasks in the forward-looking direction more difficult. We have collected underwater forward-looking stereo-vision and visual-inertial image sets in the Mediterranean and Red Sea. To our knowledge there are no other public datasets in the underwater environment acquired with this camera-sensor orientation published with ground-truth. These datasets are critical for the development of several underwater applications, including obstacle avoidance, visual odometry, 3D tracking, Simultaneous Localization and Mapping (SLAM) and depth estimation. The stereo datasets include synchronized stereo images in dynamic underwater environments with objects of known-size. The visual-inertial datasets contain monocular images and IMU measurements, aligned with millisecond resolution timestamps and objects of known size which were placed in the scene. Both sensor configurations allow for scale estimation, with the calibrated baseline in the stereo setup and the IMU in the visual-inertial setup. Ground truth depth maps were created offline for both dataset types using photogrammetry. The ground truth is validated with multiple known measurements placed throughout the imaged environment. There are 5 stereo and 8 visual-inertial datasets in total, each containing thousands of images, with a range of different underwater visibility and ambient light conditions, natural and man-made structures and dynamic camera motions. The forward-looking orientation of the camera makes these datasets unique and ideal for testing underwater obstacle-avoidance algorithms and for navigation close to the seafloor in dynamic environments. With our datasets, we hope to encourage the advancement of autonomous functionality for underwater vehicles in dynamic and/or shallow water environments.
研究动机与目标
- 通过提供公开可获取的数据集,推动水下前向观测感知与导航系统的发展。
- 提供带地面真实深度和尺度信息的同步立体与单目视觉-惯性数据,用于度量重建。
- 使在动态水下环境中评估 VI-SLAM、SLAM 和单目深度估计算法成为可能。
- 在不同能见度、照明条件和结构场景下提供数据,以对相关感知算法进行压力测试。
- 使用 Agisoft Metashape 推导的深度图及已知尺寸的标定目标进行地面真值验证。
提出的方法
- 使用了两种成像平台:由潜水员手持的立体相机与 BlueROV2 视觉-惯性系统。
- 地面真实深度图通过离线使用 Agisoft Metashape 生成,以已知尺寸的对象作为尺度参考。
- 标定程序建立了相机内外参数以及传感器变换(立体基线和 IMU-to-camera)。
- 潜水覆盖地中海与红海,具有多样的能见度、照明和运动,产生了5个立体数据集和8个视觉-惯性数据集,总计数千张图像。
- 数据集包括原始图像和 SeaErra-enhanced 图像,具有同步时间戳以及地面真值相机位姿和深度图。
实验结果
研究问题
- RQ1前向观测的水下立体视觉和视觉-惯性数据是否能够在低能见度、动态水下环境中支撑鲁棒的 SLAM 和 VO?
- RQ2通过摄影测量获得的地面真实深度与单目和立体 underwater 感知中的学习深度相比如何?
- RQ3水下成像现象(caustics、attenuation、turbidity)对这些数据集中 3D 重建精度的影响是什么?
- RQ4这些数据集中富含回环的序列对于在水下评估 VI-SLAM 和立体 SLAM 方法是否有益?
主要发现
- FLSea 集合包含在地中海和红海收集的 12 个视觉-惯性数据集和 4 个水下前向立体数据集。
- 地面真值深度图和相机位姿,离线使用 Agisoft Metashape 生成,并用已知尺寸的对象进行验证。
- 所有数据集均包含内外参数标定和通过经过标定基线(立体)或 IMU(视觉-惯性)得到的尺度信息。
- 立体数据提供以 10 Hz 同步的图像对,并配有已知尺寸的对象以进行尺度与深度验证。
- 视觉-惯性数据提供以 10 Hz 的单目图像,IMU 数据频率在 20–100 Hz,时间戳达到毫秒级别,用于尺度感知的 VIO/VI-SLAM 评估。
- 地面真值深度精度报告显示,在可验证对象上误差低于 0.5 cm。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。