Skip to main content
QUICK REVIEW

[Paper Review] D$^2$-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios

Zhengping Che, Max Guangyu Li|arXiv (Cornell University)|Apr 3, 2019
Advanced Neural Network Applications18 references39 citations
TL;DR

D2-City provides over 10,000 dashcam videos from China with dense 12-class object detection and tracking annotations on 1,000 videos, plus keyframe annotations for the rest, enabling large-scale detection, tracking, and interpolation tasks.

ABSTRACT

Driving datasets accelerate the development of intelligent driving and related computer vision technologies, while substantial and detailed annotations serve as fuels and powers to boost the efficacy of such datasets to improve learning-based models. We propose D$^2$-City, a large-scale comprehensive collection of dashcam videos collected by vehicles on DiDi's platform. D$^2$-City contains more than 10000 video clips which deeply reflect the diversity and complexity of real-world traffic scenarios in China. We also provide bounding boxes and tracking annotations of 12 classes of objects in all frames of 1000 videos and detection annotations on keyframes for the remainder of the videos. Compared with existing datasets, D$^2$-City features data in varying weather, road, and traffic conditions and a huge amount of elaborate detection and tracking annotations. By bringing a diverse set of challenging cases to the community, we expect the D$^2$-City dataset will advance the perception and related areas of intelligent driving.

Motivation & Objective

  • Provide a large-scale, diverse dashcam video dataset reflecting real-world Chinese traffic scenarios.
  • Offer dense bounding box and tracking annotations for 12 road-object classes across 1,000 videos.
  • Enable benchmarking for object detection, multi-object tracking, and large-scale detection interpolation in driving contexts.

Proposed method

  • Collect over 11,211 dashcam videos from DiDi platform across five Chinese cities.
  • Annotate bounding boxes and tracking IDs for 12 classes on all frames of 1,000 videos; provide keyframe detections on remaining videos.
  • Use CVAT-based annotation platform with frame propagation and mean-shift interpolation to balance quality and efficiency.
  • Blur license plates and faces for privacy; blur timestamps; ensure information security and policy compliance.
  • Split 1,000 annotated videos into training (700), validation (100), and test (200) sets; release training/validation annotations publicly.

Experimental results

Research questions

  • RQ1How does the D2-City dataset support robust detection and tracking under diverse weather, road, and traffic conditions in China?
  • RQ2What are the statistics (counts, occlusions, truncations) of objects and bounding boxes across the dataset?
  • RQ3Can the dataset enable large-scale detection interpolation by providing many keyframe annotations in addition to dense per-frame labels?
  • RQ4What is the distribution of road types, traffic modes, and ego-vehicle behaviors across the collected videos?

Key findings

  • The dataset contains 11,211 driving videos totaling about 100 hours, collected from roughly 500 vehicles in 5 Chinese cities.
  • For 1,000 videos (over 700,000 frames), 12 object classes are densely annotated with bounding boxes and tracking IDs; remaining videos have keyframe detections for interpolation tasks.
  • The collection includes diverse road types and conditions, with urban and suburban footage, varying speeds, and frequent intersections (average 0.26 intersections per 30s clip).
  • Average scene statistics show about 5.37 cars and 0.85 persons per frame; 45.23% of objects are occluded and 5.71% are truncated.
  • Bounding box analysis across resolutions (720p and 1080p) provides mean/median object sizes for all classes; tracking annotations indicate substantial object counts per video (e.g., 33.48 cars, 8.46 persons per video).
  • The dataset emphasizes three-wheel vehicles (open/closed tricycles) and includes a group_id mechanism to link riders with their vehicle when applicable.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.