[論文レビュー] FLSea: Underwater Visual-Inertial and Stereo-Vision Forward-Looking Datasets
FLSeaは、公開された前方視点水中ステレオおよび視覚-慣性データセットを、地上真値深度マップと較正済みセンサデータとともに提供し、困難な水中環境でのSLAM、VO、深度推定研究を可能にします。
Visibility underwater is challenging, and degrades as the distance between the subject and camera increases, making vision tasks in the forward-looking direction more difficult. We have collected underwater forward-looking stereo-vision and visual-inertial image sets in the Mediterranean and Red Sea. To our knowledge there are no other public datasets in the underwater environment acquired with this camera-sensor orientation published with ground-truth. These datasets are critical for the development of several underwater applications, including obstacle avoidance, visual odometry, 3D tracking, Simultaneous Localization and Mapping (SLAM) and depth estimation. The stereo datasets include synchronized stereo images in dynamic underwater environments with objects of known-size. The visual-inertial datasets contain monocular images and IMU measurements, aligned with millisecond resolution timestamps and objects of known size which were placed in the scene. Both sensor configurations allow for scale estimation, with the calibrated baseline in the stereo setup and the IMU in the visual-inertial setup. Ground truth depth maps were created offline for both dataset types using photogrammetry. The ground truth is validated with multiple known measurements placed throughout the imaged environment. There are 5 stereo and 8 visual-inertial datasets in total, each containing thousands of images, with a range of different underwater visibility and ambient light conditions, natural and man-made structures and dynamic camera motions. The forward-looking orientation of the camera makes these datasets unique and ideal for testing underwater obstacle-avoidance algorithms and for navigation close to the seafloor in dynamic environments. With our datasets, we hope to encourage the advancement of autonomous functionality for underwater vehicles in dynamic and/or shallow water environments.
研究の動機と目的
- 公開アクセス可能なデータセットを提供することにより、水中の前方視認・ナビゲーションシステムの開発を促進する。
- 実測スケールを含む地上真値深度とスケールを提供した、同期されたステレオおよびモノクラル視覚-慣性データを提供する。
- 動的な水中環境でのVI-SLAM、SLAM、モノクラ深度推定アルゴリズムの評価を可能にする。
- さまざまな視界、照明、構造シーンにわたるデータを提供し、関連知覚アルゴリズムを負荷試験する。
- 既知サイズの較正ターゲットを用いたAgisoft Metashape由来の深度マップによる地上真値検証。
提案手法
- 2つの撮影プラットフォームを使用: ダイバーが保持するステレオリグとBlueROV2視覚-慣性システム。
- 地上真値深度マップは、既知サイズの物体をスケール参照として用い、Agisoft Metashapeを用いてオフラインで生成された。
- 較正手順は、内部・外部キャリブレーションパラメータとセンサ変換(ステレオベースラインとIMU-to-camera)を確立した。
- ダイブは地中海と紅海をカバーし、多様な視界・照明・運動条件のもと、5つのステレオデータセットと8つの視覚-慣性データセットを得て、数千枚の画像に達した。
- データセットにはオリジナル画像とSeaErra強化画像が含まれ、同期されたタイムスタンプと地上真値のカメラ姿勢・深度マップを備える。
実験結果
リサーチクエスチョン
- RQ1低視認性で動的な水中環境において、前方視点水中ステレオと視覚-慣性データは堅牢なSLAMとVOをサポートできるか?
- RQ2フォトグラメトリによって得られた地上真値深度は、水中のモノクラルおよびステレオ知覚における学習深度とどう比較されるか?
- RQ3水中イメージング現象(caustics(光の焦点化現象))、減衰、濁度)が、これらのデータセットにおける3D再構成の精度に与える影響は?
- RQ4これらのデータセットにおけるループクロージャが豊富なシーケンスは、水中でのVI-SLAMおよびステレオSLAM手法の評価に有益か?
主な発見
- FLSeaコレクションは、地中海と紅海で収集された、水中前方視の視覚-慣性データセット12件とステレオデータセット4件からなる。
- 地上真値深度マップとカメラ姿勢が提供され、Agisoft Metashapeを用いてオフライン生成され、既知サイズの物体で検証された。
- すべてのデータセットには、内部/外部キャリブレーションと、ステレオのような較正済みベースラインまたは視覚-慣性のIMUによるスケール情報が含まれる。
- ステレオデータは、スケールと深度検証のため、10 Hzの同期画像ペアを提供する。
- 視覚-慣性データは、10 Hzのモノ画像と20–100 HzのIMUデータ、ミリ秒レベルのタイムスタンプを提供し、スケール認識VIO/VI-SLAM評価を可能にする。
- 地上真値深度の精度レポートは、検証が可能だった測定対象物で0.5 cm未満の誤差を示している。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。