[Paper Review] MM-Fi: Multi-Modal Non-Intrusive 4D Human Dataset for Versatile Wireless Sensing
MM-Fi introduces a first five-modality, non-intrusive 4D human dataset (RGB, depth, LiDAR, mmWave radar, WiFi CSI) with rich pose/action annotations on 40 subjects across 27 actions, plus baseline benchmarks for multi- and single-modal wireless sensing.
4D human perception plays an essential role in a myriad of applications, such as home automation and metaverse avatar simulation. However, existing solutions which mainly rely on cameras and wearable devices are either privacy intrusive or inconvenient to use. To address these issues, wireless sensing has emerged as a promising alternative, leveraging LiDAR, mmWave radar, and WiFi signals for device-free human sensing. In this paper, we propose MM-Fi, the first multi-modal non-intrusive 4D human dataset with 27 daily or rehabilitation action categories, to bridge the gap between wireless sensing and high-level human perception tasks. MM-Fi consists of over 320k synchronized frames of five modalities from 40 human subjects. Various annotations are provided to support potential sensing tasks, e.g., human pose estimation and action recognition. Extensive experiments have been conducted to compare the sensing capacity of each or several modalities in terms of multiple tasks. We envision that MM-Fi can contribute to wireless sensing research with respect to action recognition, human pose estimation, multi-modal learning, cross-modal supervision, and interdisciplinary healthcare research.
Motivation & Objective
- Address privacy and convenience limitations of cameras and wearables by using non-intrusive wireless sensors (LiDAR, mmWave, WiFi).
- Create a large, multi-modal 4D human dataset with extensive annotations for pose, 3D position, and actions.
- Enable multi-modal learning, cross-modal supervision, and domain generalization in wireless sensing.
- Provide benchmarks and tooling to advance research in 3D HPE and action recognition across modalities.
Proposed method
- Develop a synchronized, mobile sensor platform capturing RGB-D, LiDAR, mmWave radar, and WiFi CSI data via ROS, with a unified 10 Hz frame rate.
- Annotate 2D/3D pose, 3D body landmarks, 3D dense pose, action categories, and 3D subject position; refine 3D keypoints via optimization (L_G and L_A) on multi-view triangulated data.
- Fuse LiDAR and camera data to produce a circumscribed 3D position cube and annotate using a high-quality ground truth within ~50 mm error.
- Provide 3D dense pose labels derived from RGB-based dense pose models to enable wireless dense pose estimation experiments.
- Offer temporal action segmentation labels and a PyTorch data loader for convenient multi- and single-modal experiments.
Experimental results
Research questions
- RQ1How do five non-intrusive modalities compare for 3D human pose estimation (HPE) under various data splits and protocols?
- RQ2Can multi-modal fusion improve robustness and accuracy of 3D HPE and action recognition in wireless sensing?
- RQ3How do cross-subject and cross-environment generalization affect modality performance in MM-Fi?
- RQ4What is the feasibility and quality of wireless dense pose and action segmentation derived from multi-modal data?
Key findings
- Single-modality results show LiDAR achieving MPJPE of 98.1±2.2, 110.1±2.9, 192.3±30.4 mm and PA-MPJPE of 65.2±0.7, 66.2±1.2, 100.4±5.4 mm across P1, P2, P3 respectively (S1).
- mmWave radar yields MPJPE of 109.8±2.7, 128.4±6.9, 166.2±4.5 mm and PA-MPJPE of 55.6±1.4, 58.7±4.3, 73.9±2.7 mm (S1).
- WiFi CSI-based 3D HPE under random splits shows MPJPE around 367.8±0.9 to 369.5±0.3 mm and PA-MPJPE around 121.0±2.2 to 121.4±0.1 mm (S3).
- Cross-subject results indicate LiDAR and mmWave generalize well (PA-MPJPE changes within a few mm), while WiFi shows degraded generalization due to limited resolution.
- Multi-modal fusion (e.g., RGB+LiDAR or R+L+W) improves HPE performance over single modalities in several settings, with I+L and R+L+W achieving notable gains in MPJPE/PA-MPJPE across protocols.
- In cross-environment scenarios, mmWave-based 3D HPE remains the most robust among modalities, with LiDAR and WiFi experiencing larger performance drops; fusion can mitigate some losses.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.