Skip to main content
QUICK REVIEW

[Paper Review] TICaM: A Time-of-flight In-car Cabin Monitoring Dataset

Jigyasa Katrolia, Bruno Mirbach|arXiv (Cornell University)|Mar 22, 2021
Video Surveillance and Tracking Methods15 references16 citations
TL;DR

TICaM is a large-scale, multi-modal in-car cabin monitoring dataset featuring 6.7K real and 3.3K synthetic time-of-flight depth, RGB, and infrared images with comprehensive annotations for 2D/3D detection, instance and semantic segmentation, and activity recognition. It uniquely combines real and synthetic data to enable robust training and domain adaptation evaluation for autonomous vehicle safety and human-vehicle interaction systems.

ABSTRACT

We present TICaM, a Time-of-flight In-car Cabin Monitoring dataset for vehicle interior monitoring using a single wide-angle depth camera. Our dataset addresses the deficiencies of currently available in-car cabin datasets in terms of the ambit of labeled classes, recorded scenarios and provided annotations; all at the same time. We record an exhaustive list of actions performed while driving and provide for them multi-modal labeled images (depth, RGB and IR), with complete annotations for 2D and 3D object detection, instance and semantic segmentation as well as activity annotations for RGB frames. Additional to real recordings, we provide a synthetic dataset of in-car cabin images with same multi-modality of images and annotations, providing a unique and extremely beneficial combination of synthetic and real data for effectively training cabin monitoring systems and evaluating domain adaptation approaches. The dataset is available at https://vizta-tof.kl.dfki.de/.

Motivation & Objective

  • Address the lack of comprehensive in-car cabin datasets that cover diverse occupant types, real-world scenarios, and multi-modal annotations.
  • Overcome limitations in existing datasets by including underrepresented scenarios such as passengers, children in forward/rear-facing seats, and everyday objects.
  • Provide a unified dataset with synchronized depth, RGB, and infrared data for multi-modal computer vision tasks in automotive environments.
  • Enable domain adaptation research by combining high-quality real data with realistic synthetic data generated using Blender.
  • Support the development of safety-critical systems such as adaptive airbag deployment and driver distraction monitoring through detailed, multi-task annotations.

Proposed method

  • Record real in-car cabin scenes using a Kinect Azure mounted near the rear-view mirror to capture time-of-flight depth, RGB, and infrared images.
  • Capture 123K RGB frames with synchronized activity annotations for driver and passenger behaviors across 20 distinct activity classes.
  • Manually annotate real data using the SALT 3D-annotation tool for 2D and 3D bounding boxes, instance and class segmentation masks, and activity labels.
  • Generate 3.3K synthetic images using Blender with automatically generated ground truth for 2D/3D detection, segmentation, and activity labels.
  • Ensure alignment between depth and RGB cameras by calibrating and providing known extrinsic parameters for cross-modal training.
  • Split real data into training (86K RGB frames) and testing (37K RGB frames) sets, using synthetic data only for training to evaluate domain adaptation.

Experimental results

Research questions

  • RQ1Can a combined real and synthetic time-of-flight in-car cabin dataset improve the robustness and generalization of cabin monitoring systems?
  • RQ2To what extent does the inclusion of diverse scenarios—such as passengers, children in rear-facing seats, and common objects—enhance the practical utility of in-car monitoring datasets?
  • RQ3How effective is synthetic data in enabling domain adaptation for real-world in-car depth-based perception tasks?
  • RQ4Can multi-modal annotations (depth, RGB, IR, 2D/3D detection, segmentation, activity) jointly improve performance in driver state estimation and safety applications?
  • RQ5Does the availability of synchronized, multi-resolution annotations across depth, RGB, and infrared modalities reduce annotation cost and improve model training efficiency?

Key findings

  • TICaM provides 6.7K real time-of-flight depth images and 123K RGB frames with activity annotations, covering 20 distinct activity classes including complex behaviors like turning while looking left or right.
  • The dataset includes 3.3K synthetic images with identical annotation formats to real data, enabling direct evaluation of domain adaptation techniques.
  • The real data contains 4.7K images for 2D and 3D detection training and 2.0K for testing, with full 3D bounding boxes, segmentation masks, and activity labels.
  • The synthetic data is generated with Blender using accurate camera models and object placement, ensuring high visual and geometric fidelity to real data.
  • The dataset supports multi-task learning with synchronized annotations across depth, RGB, and infrared modalities, enabling cross-modal model training.
  • The inclusion of low-remission flags for reflective or dark objects improves detection robustness in challenging lighting and material conditions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.