[Paper Review] S3E: A Multi-Robot Multimodal Dataset for Collaborative SLAM
S3E introduces a large-scale, multimodal dataset for collaborative SLAM, recorded by three unmanned ground vehicles equipped with LiDAR, stereo cameras, and IMUs across 12 diverse outdoor and indoor sequences. The dataset enables evaluation of collaborative SLAM systems under realistic inter-robot loop closure scenarios, revealing that state-of-the-art methods still struggle with low-overlap and degenerate environments, highlighting a critical need for improved robustness in real-world multi-robot systems.
The burgeoning demand for collaborative robotic systems to execute complex tasks collectively has intensified the research community's focus on advancing simultaneous localization and mapping (SLAM) in a cooperative context. Despite this interest, the scalability and diversity of existing datasets for collaborative trajectories remain limited, especially in scenarios with constrained perspectives where the generalization capabilities of Collaborative SLAM (C-SLAM) are critical for the feasibility of multi-agent missions. Addressing this gap, we introduce S3E, an expansive multimodal dataset. Captured by a fleet of unmanned ground vehicles traversing four distinct collaborative trajectory paradigms, S3E encompasses 13 outdoor and 5 indoor sequences. These sequences feature meticulously synchronized and spatially calibrated data streams, including 360-degree LiDAR point cloud, high-resolution stereo imagery, high-frequency inertial measurement units (IMU), and Ultra-wideband (UWB) relative observations. Our dataset not only surpasses previous efforts in scale, scene diversity, and data intricacy but also provides a thorough analysis and benchmarks for both collaborative and individual SLAM methodologies. For access to the dataset and the latest information, please visit our repository at https://pengyu-team.github.io/S3E.
Motivation & Objective
- Address the lack of large-scale, realistic, multimodal datasets for collaborative SLAM (C-SLAM) that support systematic benchmarking and reproducibility.
- Overcome the limitations of existing datasets by capturing long-duration, temporally synchronized, and spatially calibrated data from multiple agents in diverse indoor and outdoor environments.
- Design four distinct collaborative trajectory paradigms to systematically evaluate inter-robot loop closure performance under varying degrees of spatial and temporal overlap.
- Provide ground truth localization using RTK GNSS and motion capture systems to enable accurate evaluation of state estimation accuracy.
- Establish baselines for both single-agent and collaborative SLAM using state-of-the-art algorithms across multiple sensor modalities.
Proposed method
- Deployed a fleet of three remote-controlled unmanned ground vehicles (UGVs) equipped with 16-beam LiDAR, stereo cameras, 9-axis IMUs, and dual-antenna RTK GNSS for high-precision localization.
- Recorded 12 sequences—7 outdoor and 5 indoor—each exceeding 200 seconds, with full temporal synchronization and spatial calibration across all sensors.
- Designed four collaborative trajectory paradigms to induce different levels of intra- and inter-robot loop closure opportunities, including concentric circles, intersection curves, open-set degenerate environments, and co-directional movement with limited overlap.
- Generated ground truth using dual-antenna RTK GNSS in GNSS-available areas and motion capture systems for indoor sequences, ensuring centimeter-level accuracy.
- Applied similarity transformation-based ATE (Absolute Trajectory Error) computation to evaluate trajectory accuracy, using least squares estimation to align predicted and ground truth trajectories.
- Evaluated state-of-the-art C-SLAM and single-agent SLAM methods (e.g., LIO-SAM, DCL-SLAM, COVINS, DiSCo-SLAM, LVI-SAM, ORB-SLAM3) across different sensor modalities and trajectory types.
Experimental results
Research questions
- RQ1How do existing collaborative SLAM methods perform under realistic, large-scale, multimodal, and long-duration multi-robot scenarios with varying degrees of inter-robot trajectory overlap?
- RQ2To what extent does collaborative SLAM improve localization accuracy and robustness compared to single-agent SLAM in challenging environments with limited visual or LiDAR overlap?
- RQ3What are the key failure modes of current C-SLAM systems when inter-robot loop closures are sparse or occur in degenerate geometric configurations?
- RQ4How does sensor modality (LiDAR, visual, inertial) influence the performance and reliability of collaborative SLAM in real-world deployments?
- RQ5What role does trajectory design—particularly temporal and spatial diversity—play in enabling effective inter-robot loop closure and back-end optimization?
Key findings
- The S3E dataset contains 12 sequences (7 outdoor, 5 indoor), each lasting over 200 seconds, with 4× longer average recording time than the pioneering EuRoC dataset.
- LiDAR-based SLAM methods consistently outperformed visual-based methods in long-term tracking, especially in cornering or low-texture scenarios where visual features were lost.
- DCL-SLAM reduced average ATE by 0.42 compared to its single-agent counterpart (LIO-SAM), demonstrating significant improvement through inter-robot loop closure detection.
- In the Playground_1 sequence, both Alpha and Bob failed to track frames with single-agent LIO-SAM due to cornering, but succeeded when collaborating via DCL-SLAM, highlighting the robustness gain from collaboration.
- COVINS reduced average ATE by 7.09 compared to its single-agent version in the Square_1 sequence, and achieved 1.75 ATE with collaboration, while ORB-SLAM3 failed to track in the same scenario.
- DCL-SLAM outperformed COVINS in low-overlap scenarios, particularly at endpoints with co-directional movement, due to more effective feature matching and loop closure detection under sparse overlap.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.