Skip to main content
QUICK REVIEW

[论文解读] DETR for Crowd Pedestrian Detection

Matthieu Lin, Chuming Li|arXiv (Cornell University)|Dec 12, 2020
Advanced Neural Network Applications参考文献 39被引用 38
一句话总结

PED-DETR 是一个行人端到端检测器,分析为何 DETR 在人群场景中表现不佳,并引入 DQRF、V-Match 和 Fast-KM,在 CrowdHuman 和 CityPersons 上实现最先进的结果。

ABSTRACT

Pedestrian detection in crowd scenes poses a challenging problem due to the heuristic defined mapping from anchors to pedestrians and the conflict between NMS and highly overlapped pedestrians. The recently proposed end-to-end detectors(ED), DETR and deformable DETR, replace hand designed components such as NMS and anchors using the transformer architecture, which gets rid of duplicate predictions by computing all pairwise interactions between queries. Inspired by these works, we explore their performance on crowd pedestrian detection. Surprisingly, compared to Faster-RCNN with FPN, the results are opposite to those obtained on COCO. Furthermore, the bipartite match of ED harms the training efficiency due to the large ground truth number in crowd scenes. In this work, we identify the underlying motives driving ED's poor performance and propose a new decoder to address them. Moreover, we design a mechanism to leverage the less occluded visible parts of pedestrian specifically for ED, and achieve further improvements. A faster bipartite match algorithm is also introduced to make ED training on crowd dataset more practical. The proposed detector PED(Pedestrian End-to-end Detector) outperforms both previous EDs and the baseline Faster-RCNN on CityPersons and CrowdHuman. It also achieves comparable performance with state-of-the-art pedestrian detection methods. Code will be released soon.

研究动机与目标

  • 评估为何 DETR 和 deformable DETR 在拥挤人群检测中的表现不如 Faster-RCNN with FPN。
  • 提出 Dense Query and Rectified Attention Field decoder (DQRF) 以提升 DETR 在行人检测上的性能。
  • 利用 V-Match 对可见区域标注来提升基于 DETR 的行人检测。
  • 引入更快的二分匹配 (Fast-KM),以使在拥挤数据集上训练 DETR 成为可行。
  • 在 CrowdHuman 和 CityPersons 基准上展示最先进或具有竞争力的性能。

提出的方法

  • 分析在拥挤人群场景中 DETR 解码器的行为,以识别失败模式。
  • 开发 DQRF 解码器,使其能够实现密集查询和一个经过整正、覆盖更广的注意力场。
  • 引入 RF(Rectified Attention Field)以在解码器层之间稳定跨注意力。
  • 提出 V-Match,以在各层对完整的和可见的行人区域进行监督且不产生额外成本。
  • 实现 Fast-KM,以在训练期间加速匈牙利匹配步骤。

实验结果

研究问题

  • RQ1为何原始的 DETR 和 deformable DETR 在拥挤人群检测中不如 Faster-RCNN?
  • RQ2Dense Query and Rectified Attention Field 解码器是否能够提升 DETR 对拥挤行人检测的效果?
  • RQ3通过 V-Match 利用可见区域标注是否能在不增加额外成本的情况下提升端到端行人检测?
  • RQ4在不损失精度的前提下,Fast-KM 能在匈牙利匹配训练中实现多少加速?

主要发现

  • PED 在 CrowdHuman 和 CityPersons 上相对于 deformable DETR 和 Faster-RCNN 基线取得了改进。
  • 所提出的 DQRF 解码器显著缩小了在拥挤场景中基于 DETR 的检测器与 Faster-RCNN 之间的差距。
  • V-Match 提供了可见区域监督的增益且没有额外成本。
  • Fast-KM 在训练期间的二分匹配可实现高达 10x 的加速。
  • PED 在具有挑战性的基准测试中达到与最先进方法竞争的结果。
  • 密集查询和整正的注意力场在拥挤且遮挡的场景中减少了漏检。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。