[Paper Review] FLSea: Underwater Visual-Inertial and Stereo-Vision Forward-Looking Datasets
FLSea presents public forward-looking underwater stereo and visual-inertial datasets with ground-truth depth maps and calibrated sensor data to enable SLAM, VO, and depth estimation research in challenging underwater environments.
Visibility underwater is challenging, and degrades as the distance between the subject and camera increases, making vision tasks in the forward-looking direction more difficult. We have collected underwater forward-looking stereo-vision and visual-inertial image sets in the Mediterranean and Red Sea. To our knowledge there are no other public datasets in the underwater environment acquired with this camera-sensor orientation published with ground-truth. These datasets are critical for the development of several underwater applications, including obstacle avoidance, visual odometry, 3D tracking, Simultaneous Localization and Mapping (SLAM) and depth estimation. The stereo datasets include synchronized stereo images in dynamic underwater environments with objects of known-size. The visual-inertial datasets contain monocular images and IMU measurements, aligned with millisecond resolution timestamps and objects of known size which were placed in the scene. Both sensor configurations allow for scale estimation, with the calibrated baseline in the stereo setup and the IMU in the visual-inertial setup. Ground truth depth maps were created offline for both dataset types using photogrammetry. The ground truth is validated with multiple known measurements placed throughout the imaged environment. There are 5 stereo and 8 visual-inertial datasets in total, each containing thousands of images, with a range of different underwater visibility and ambient light conditions, natural and man-made structures and dynamic camera motions. The forward-looking orientation of the camera makes these datasets unique and ideal for testing underwater obstacle-avoidance algorithms and for navigation close to the seafloor in dynamic environments. With our datasets, we hope to encourage the advancement of autonomous functionality for underwater vehicles in dynamic and/or shallow water environments.
Motivation & Objective
- Motivate development of underwater forward-looking perception and navigation systems by providing publicly accessible datasets.
- Provide synchronized stereo and monocular visual-inertial data with ground-truth depth and scale for metric reconstruction.
- Enable evaluation of VI-SLAM, SLAM, and monocular depth estimation algorithms in dynamic underwater environments.
- Offer data across varied visibility, lighting, and structural scenarios to stress-test related perception algorithms.
- Ground-truth validation using Agisoft Metashape-derived depth maps with known-size calibration targets.
Proposed method
- Two imaging platforms were used: a diver-held stereo rig and a BlueROV2 visual-inertial system.
- Ground-truth depth maps were generated offline using Agisoft Metashape with objects of known size as scale references.
- Calibration procedures established intrinsic and extrinsic camera parameters and sensor transformations (stereo baseline and IMU-to-camera).
- Dives covered the Mediterranean and Red Sea with diverse visibility, lighting, and motion, yielding 5 stereo and 8 visual-inertial datasets totaling thousands of images.
- Datasets include original and SeaErra-enhanced images, with synchronized timestamps and ground-truth camera poses and depth maps.
Experimental results
Research questions
- RQ1Can forward-looking underwater stereo and visual-inertial data support robust SLAM and VO in low-visibility, dynamic underwater environments?
- RQ2How does ground-truth depth obtained via photogrammetry compare to learned depth in monocular and stereo underwater perception?
- RQ3What are the effects of underwater imaging phenomena (caustics, attenuation, turbidity) on 3D reconstruction accuracy in these datasets?
- RQ4Are loop-closure-rich sequences in these datasets beneficial for evaluating VI-SLAM and stereo SLAM methods under water?
Key findings
- The FLSea collection comprises 12 visual-inertial and 4 stereo underwater forward-looking datasets collected in the Mediterranean and Red Sea.
- Ground truth depth maps and camera poses are provided, generated offline with Agisoft Metashape and validated with known-size objects.
- All datasets include intrinsic/extrinsic calibrations and scale information via calibrated baselines (stereo) or IMU (visual-inertial).
- The stereo data provide synchronized image pairs at 10 Hz with objects of known size for scale and depth validation.
- The visual-inertial data provide monocular images at 10 Hz with IMU data at 20–100 Hz and millisecond-level timestamps for scale-aware VIO/VI-SLAM evaluation.
- Ground-truth depth accuracy reports indicate sub-0.5 cm errors for measured objects where validation was possible.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.