[Paper Review] The LuViRA Dataset: Synchronized Vision, Radio, and Audio Sensors for Indoor Localization
The LuViRA dataset provides synchronized, high-precision measurements from vision, radio, and audio sensors in an indoor environment, enabling research on multimodal sensor fusion for centimeter-level indoor localization. It includes 88 trajectories with 0.5 mm 6DoF ground truth, 12-microphone audio, RGB/depth images, IMU data, and massive MIMO channel responses, all time-synchronized for robust algorithm evaluation and development.
We present a synchronized multisensory dataset for accurate and robust indoor localization: the Lund University Vision, Radio, and Audio (LuViRA) Dataset. The dataset includes color images, corresponding depth maps, inertial measurement unit (IMU) readings, channel response between a 5G massive multiple-input and multiple-output (MIMO) testbed and user equipment, audio recorded by 12 microphones, and accurate six degrees of freedom (6DOF) pose ground truth of 0.5 mm. We synchronize these sensors to ensure that all data is recorded simultaneously. A camera, speaker, and transmit antenna are placed on top of a slowly moving service robot, and 89 trajectories are recorded. Each trajectory includes 20 to 50 seconds of recorded sensor data and ground truth labels. Data from different sensors can be used separately or jointly to perform localization tasks, and data from the motion capture (mocap) system is used to verify the results obtained by the localization algorithms. The main aim of this dataset is to enable research on sensor fusion with the most commonly used sensors for localization tasks. Moreover, the full dataset or some parts of it can also be used for other research areas such as channel estimation, image classification, etc. Our dataset is available at: https://github.com/ilaydayaman/LuViRA_Dataset
Motivation & Objective
- Address the lack of public, synchronized datasets combining vision, radio, and audio sensors for indoor localization research.
- Enable development and evaluation of multimodal localization algorithms that improve accuracy, reliability, and energy efficiency.
- Provide a benchmark for fusing data from cameras, RF modules, and microphones to overcome individual sensor limitations.
- Support research beyond localization, such as channel estimation, audio source localization, and image classification.
- Facilitate real-time, centimeter-level localization in dynamic indoor environments like smart factories and autonomous service robots.
Proposed method
- Deployed a mobile industrial robot (MIR200) equipped with a camera, speaker, and transmit antenna in a motion capture studio to record 88 synchronized trajectories.
- Collected RGB images, depth maps, IMU readings, 12-microphone audio, and massive MIMO channel state information (CSI) with 0.5 mm 6DoF ground truth from a motion capture system.
- Used a time synchronization unit to align all sensor data across vision, audio, radio, and inertial systems with sub-millisecond precision.
- Recorded data from 1 second before to 1 second after each trajectory, ensuring full temporal coverage of movement.
- Provided background noise recordings to support noise suppression in audio-based localization algorithms.
- Calibrated all sensor systems and accounted for environmental factors such as temperature and cable interference in data interpretation.
Experimental results
Research questions
- RQ1How does fusing vision, radio, and audio sensor data improve localization accuracy and robustness in dynamic indoor environments?
- RQ2To what extent can synchronized multimodal data reduce localization latency and power consumption in real-time applications?
- RQ3How do environmental factors like temperature changes and moving obstacles affect the performance of vision, audio, and radio-based localization systems?
- RQ4What is the impact of cable interference and line-of-sight blockages on radio and vision-based localization in real-world deployments?
- RQ5Can the LuViRA dataset serve as a reliable benchmark for training and evaluating AI/ML models in 6G and low-power localization systems?
Key findings
- The LuViRA dataset provides 88 synchronized indoor trajectories with 0.5 mm 6DoF ground truth, enabling centimeter-level localization accuracy.
- Vision-based localization is highly accurate in well-lit, static conditions but fails in darkness; audio and radio systems remain functional under low-light conditions.
- Audio-based localization is significantly degraded in high-noise environments, especially when people move around, leading to frame drops in ground truth.
- Radio-based localization is affected by blockages caused by people and cables, particularly in dynamic trajectories with human movement.
- The dataset includes background noise recordings to support noise cancellation in audio-based localization algorithms, improving robustness in noisy environments.
- Temperature increases in the motion capture studio (up to 28 °C) may affect radio and audio propagation, particularly impacting the 'grid' trajectories with consistent viewing and transmission directions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.