[Paper Review] ConfLab: A Data Collection Concept, Dataset, and Benchmark for Machine Analysis of Free-Standing Social Interactions in the Wild
ConfLab introduces a novel, privacy-preserving, multimodal data collection framework for in-the-wild social interactions, using overhead cameras and wearable 9-axis IMUs to capture high-fidelity, synchronized data from 48 participants at a real conference. The dataset enables new benchmarks in pose estimation, speaker detection, and F-formation analysis, achieving state-of-the-art results in keypoint detection (AP75: 0.82) and speaker detection (AUC: 0.92) using minimal audio and body motion cues.
Recording the dynamics of unscripted human interactions in the wild is challenging due to the delicate trade-offs between several factors: participant privacy, ecological validity, data fidelity, and logistical overheads. To address these, following a 'datasets for the community by the community' ethos, we propose the Conference Living Lab (ConfLab): a new concept for multimodal multisensor data collection of in-the-wild free-standing social conversations. For the first instantiation of ConfLab described here, we organized a real-life professional networking event at a major international conference. Involving 48 conference attendees, the dataset captures a diverse mix of status, acquaintance, and networking motivations. Our capture setup improves upon the data fidelity of prior in-the-wild datasets while retaining privacy sensitivity: 8 videos (1920x1080, 60 fps) from a non-invasive overhead view, and custom wearable sensors with onboard recording of body motion (full 9-axis IMU), privacy-preserving low-frequency audio (1250 Hz), and Bluetooth-based proximity. Additionally, we developed custom solutions for distributed hardware synchronization at acquisition and time-efficient continuous annotation of body keypoints and actions at high sampling rates. Our benchmarks showcase some of the open research tasks related to in-the-wild privacy-preserving social data analysis: keypoints detection from overhead camera views, skeleton-based no-audio speaker detection, and F-formation detection.
Motivation & Objective
- To address the lack of high-fidelity, ecologically valid, and privacy-preserving datasets for analyzing unscripted social interactions in real-world settings.
- To enable research on fine-grained social dynamics such as F-formation, speaker turn-taking, and body gesture recognition in natural environments.
- To overcome limitations in prior datasets, including poor pose annotation, low participant count, and insufficient temporal resolution.
- To establish a community-driven data collection model that integrates ethical considerations into sensor design and data acquisition.
- To provide a benchmark for machine analysis of social behavior using minimal audio and full-body motion data.
Proposed method
- Deployed a distributed sensor setup with eight 1080p@60fps overhead cameras and custom wearable devices with 9-axis IMUs, privacy-preserving 1250 Hz audio, and Bluetooth proximity tracking.
- Implemented sub-second crossmodal synchronization (~13 ms latency) across multiple sensor streams using custom hardware synchronization protocols.
- Collected continuous, high-rate annotations of 17 full-body keypoints and social actions using time-efficient, continuous labeling pipelines.
- Used a by-the-community-for-the-community ethos to recruit 48 real conference attendees with diverse social motivations (newcomers, veterans, networking, acquaintances).
- Developed and evaluated three core benchmarks: top-down pose estimation, skeleton-based no-audio speaker detection, and F-formation detection from body poses.
- Applied advanced models including RSN for keypoint detection, Minirocket for IMU-based speaker classification, and GTCG/GCFF for F-formation analysis.
Experimental results
Research questions
- RQ1Can high-fidelity, privacy-preserving, and ecologically valid data be collected in real-world social settings without influencing natural behavior?
- RQ2How well can body keypoints be detected from overhead camera views in unconstrained, in-the-wild settings?
- RQ3To what extent can speaker status be predicted using only body motion (IMU) and low-bandwidth audio, without direct face or voice data?
- RQ4Can F-formation configurations be reliably detected from full-body pose data alone, without relying on facial or audio cues?
- RQ5How does the integration of multiple sensor modalities improve the accuracy of social behavior analysis in naturalistic environments?
Key findings
- The ConfLab dataset achieved a mean average precision (AP75) of 0.82 for 17-keypoint detection from overhead camera views, demonstrating strong performance despite the challenges of top-down perspective.
- Speaker status detection using only IMU data achieved an AUC of 0.92, showing that body motion alone can effectively identify speaking turns without audio or visual face data.
- F-formation detection using GTCG and GCFF methods achieved F1 scores above 0.75 on average across multiple cameras, proving that social interaction geometry can be inferred from pose data.
- The addition of lower-body keypoints significantly improved keypoint detection performance, with AP75 increasing from 0.75 to 0.82, indicating that leg movements contribute meaningfully to pose estimation.
- The system achieved sub-second crossmodal synchronization (~13 ms latency), enabling fine-grained temporal analysis of social interactions.
- The dataset supports complex social dynamics such as group splitting and merging, with 48 participants and diverse social motivations captured in a real-world setting.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.