[论文解读] D$^2$-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios
D2-City 提供来自中国的超过 10,000 条行车记录仪视频,在 1,000 条视频上拥有密集的 12 类对象检测与跟踪注释,其余视频提供关键帧注释,从而实现大规模检测、跟踪和插值任务。
Driving datasets accelerate the development of intelligent driving and related computer vision technologies, while substantial and detailed annotations serve as fuels and powers to boost the efficacy of such datasets to improve learning-based models. We propose D$^2$-City, a large-scale comprehensive collection of dashcam videos collected by vehicles on DiDi's platform. D$^2$-City contains more than 10000 video clips which deeply reflect the diversity and complexity of real-world traffic scenarios in China. We also provide bounding boxes and tracking annotations of 12 classes of objects in all frames of 1000 videos and detection annotations on keyframes for the remainder of the videos. Compared with existing datasets, D$^2$-City features data in varying weather, road, and traffic conditions and a huge amount of elaborate detection and tracking annotations. By bringing a diverse set of challenging cases to the community, we expect the D$^2$-City dataset will advance the perception and related areas of intelligent driving.
研究动机与目标
- 提供一个大规模、多样化的行车记录仪视频数据集,反映真实世界的中国交通场景。
- 为 1,000 条视频中的 12 个路上对象类别提供密集的边框和跟踪注释。
- 在驾驶场景中为目标检测、多对象跟踪和大规模检测插值提供基准测试。
提出的方法
- 从 DiDi 平台在五个中国城市收集超过 11,211 条行车记录仪视频。
- 对 1,000 条视频的所有帧进行 12 个类别的边框和跟踪 ID 注释;对剩余视频提供关键帧检测。
- 使用基于 CVAT 的标注平台,结合帧传播和均值漂移插值以在质量与效率之间取得平衡。
- 对车牌和人脸进行模糊处理以保护隐私;对时间戳进行模糊处理;确保信息安全和政策合规。
- 将 1,000 条带注释的视频分为训练集(700)、验证集(100)和测试集(200);公开发布训练/验证注释。
实验结果
研究问题
- RQ1D2-City 数据集如何在中国多样的天气、道路和交通条件下支持稳健的检测与跟踪?
- RQ2数据集中对象和边框的统计信息(数量、遮挡、截断)是什么?
- RQ3数据集是否能够通过提供大量关键帧注释以及密集逐帧注释来实现大规模的检测插值?
- RQ4在所收集的视频中,道路类型、交通模式和自车行为的分布是怎样的?
主要发现
- 该数据集包含 11,211 条行驶视频,总时长约 100 小时,来自中国 5 个城市约 500 辆车辆。
- 对于 1,000 条视频(超过 700,000 帧),12 个对象类别密集标注了边框和跟踪 ID;其余视频具有用于插值任务的关键帧检测。
- 收集覆盖多样的道路类型与条件,包含城市和郊区镜头、不同速度、以及频繁的交叉口(每 30 秒片段平均 0.26 个交叉口)。
- 平均场景统计显示每帧约有 5.37 辆车和 0.85 名行人;45.23% 的对象被遮挡,5.71% 被截断。
- 跨分辨率(720p 与 1080p)的边框分析提供所有类别的平均/中位对象大小;跟踪注释表明每个视频中对象数量较多(如每个视频 33.48 辆车,8.46 名行人)。
- 数据集强调三轮车辆(开式/闭式三轮车),并包含 group_id 机制以在适用时将乘员与车辆关联。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。