Skip to main content
QUICK REVIEW

[Paper Review] ConfLab: A Data Collection Concept, Dataset, and Benchmark for Machine Analysis of Free-Standing Social Interactions in the Wild

Chirag Raman, Jose Vargas-Quiros|arXiv (Cornell University)|May 10, 2022
Human Mobility and Location-Based Analysis4 citations
TL;DR

ConfLab introduces a novel, privacy-preserving, multimodal data collection framework for in-the-wild social interactions, using overhead cameras and wearable 9-axis IMUs to capture high-fidelity, synchronized data from 48 participants at a real conference. The dataset enables new benchmarks in pose estimation, speaker detection, and F-formation analysis, achieving state-of-the-art results in keypoint detection (AP75: 0.82) and speaker detection (AUC: 0.92) using minimal audio and body motion cues.

ABSTRACT

Recording the dynamics of unscripted human interactions in the wild is challenging due to the delicate trade-offs between several factors: participant privacy, ecological validity, data fidelity, and logistical overheads. To address these, following a 'datasets for the community by the community' ethos, we propose the Conference Living Lab (ConfLab): a new concept for multimodal multisensor data collection of in-the-wild free-standing social conversations. For the first instantiation of ConfLab described here, we organized a real-life professional networking event at a major international conference. Involving 48 conference attendees, the dataset captures a diverse mix of status, acquaintance, and networking motivations. Our capture setup improves upon the data fidelity of prior in-the-wild datasets while retaining privacy sensitivity: 8 videos (1920x1080, 60 fps) from a non-invasive overhead view, and custom wearable sensors with onboard recording of body motion (full 9-axis IMU), privacy-preserving low-frequency audio (1250 Hz), and Bluetooth-based proximity. Additionally, we developed custom solutions for distributed hardware synchronization at acquisition and time-efficient continuous annotation of body keypoints and actions at high sampling rates. Our benchmarks showcase some of the open research tasks related to in-the-wild privacy-preserving social data analysis: keypoints detection from overhead camera views, skeleton-based no-audio speaker detection, and F-formation detection.

Motivation & Objective

  • To address the lack of high-fidelity, ecologically valid, and privacy-preserving datasets for analyzing unscripted social interactions in real-world settings.
  • To enable research on fine-grained social dynamics such as F-formation, speaker turn-taking, and body gesture recognition in natural environments.
  • To overcome limitations in prior datasets, including poor pose annotation, low participant count, and insufficient temporal resolution.
  • To establish a community-driven data collection model that integrates ethical considerations into sensor design and data acquisition.
  • To provide a benchmark for machine analysis of social behavior using minimal audio and full-body motion data.

Proposed method

  • Deployed a distributed sensor setup with eight 1080p@60fps overhead cameras and custom wearable devices with 9-axis IMUs, privacy-preserving 1250 Hz audio, and Bluetooth proximity tracking.
  • Implemented sub-second crossmodal synchronization (~13 ms latency) across multiple sensor streams using custom hardware synchronization protocols.
  • Collected continuous, high-rate annotations of 17 full-body keypoints and social actions using time-efficient, continuous labeling pipelines.
  • Used a by-the-community-for-the-community ethos to recruit 48 real conference attendees with diverse social motivations (newcomers, veterans, networking, acquaintances).
  • Developed and evaluated three core benchmarks: top-down pose estimation, skeleton-based no-audio speaker detection, and F-formation detection from body poses.
  • Applied advanced models including RSN for keypoint detection, Minirocket for IMU-based speaker classification, and GTCG/GCFF for F-formation analysis.

Experimental results

Research questions

  • RQ1Can high-fidelity, privacy-preserving, and ecologically valid data be collected in real-world social settings without influencing natural behavior?
  • RQ2How well can body keypoints be detected from overhead camera views in unconstrained, in-the-wild settings?
  • RQ3To what extent can speaker status be predicted using only body motion (IMU) and low-bandwidth audio, without direct face or voice data?
  • RQ4Can F-formation configurations be reliably detected from full-body pose data alone, without relying on facial or audio cues?
  • RQ5How does the integration of multiple sensor modalities improve the accuracy of social behavior analysis in naturalistic environments?

Key findings

  • The ConfLab dataset achieved a mean average precision (AP75) of 0.82 for 17-keypoint detection from overhead camera views, demonstrating strong performance despite the challenges of top-down perspective.
  • Speaker status detection using only IMU data achieved an AUC of 0.92, showing that body motion alone can effectively identify speaking turns without audio or visual face data.
  • F-formation detection using GTCG and GCFF methods achieved F1 scores above 0.75 on average across multiple cameras, proving that social interaction geometry can be inferred from pose data.
  • The addition of lower-body keypoints significantly improved keypoint detection performance, with AP75 increasing from 0.75 to 0.82, indicating that leg movements contribute meaningfully to pose estimation.
  • The system achieved sub-second crossmodal synchronization (~13 ms latency), enabling fine-grained temporal analysis of social interactions.
  • The dataset supports complex social dynamics such as group splitting and merging, with 48 participants and diverse social motivations captured in a real-world setting.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.