Skip to main content
QUICK REVIEW

[논문 리뷰] Geometry-aware data augmentation for monocular 3D object detection.

Qing Lian, Botao Ye|arXiv (Cornell University)|2021. 04. 12.
Advanced Neural Network Applications참고 문헌 43인용 수 8
한 줄 요약

이 논문은 단안 3D 객체 검출에서 깊이 추정의 강인성을 향상시키기 위해 카메라 파라미터 및 객체 배치의 기하학적 이동을 시뮬레이션함으로써 기하학적 인지 데이터 증강 기법을 제안한다. 이미지 및 인스턴스 수준의 증강 과정에서 3D 기하학을 유지함으로써, KITTI 및 nuScenes 벤치마크에서 성능을 크게 향상시켜 최신 기술 수준의 성능을 달성한다.

ABSTRACT

This paper focuses on monocular 3D object detection, one of the essential modules in autonomous driving systems. A key challenge is that the depth recovery problem is ill-posed in monocular data. In this work, we first conduct a thorough analysis to reveal how existing methods fail to robustly estimate depth when different geometry shifts occur. In particular, through a series of image-based and instance-based manipulations for current detectors, we illustrate existing detectors are vulnerable in capturing the consistent relationships between depth and both object apparent sizes and positions. To alleviate this issue and improve the robustness of detectors, we convert the aforementioned manipulations into four corresponding 3D-aware data augmentation techniques. At the image-level, we randomly manipulate the camera system, including its focal length, receptive field and location, to generate new training images with geometric shifts. At the instance level, we crop the foreground objects and randomly paste them to other scenes to generate new training instances. All the proposed augmentation techniques share the virtue that geometry relationships in objects are preserved while their geometry is manipulated. In light of the proposed data augmentation methods, not only the instability of depth recovery is effectively alleviated, but also the final 3D detection performance is significantly improved. This leads to superior improvements on the KITTI and nuScenes monocular 3D detection benchmarks with state-of-the-art results.

연구 동기 및 목표

  • 단안 영상에서 정의되지 않은 깊이 추정으로 인해 발생하는 깊이 복구의 불안정성을 해결하기 위해.
  • 카메라 파라미터 및 객체 위치의 기하학적 이동에 노출되었을 때 기존 검출기의 취약점을 규명하기 위해.
  • 통제된 기하학적 변형을 도입하면서도 3D 기하학적 관계를 유지하는 데이터 증강 기법을 개발하기 위해.
  • 다양한 기하학적 변환 하에서 일반화 능력과 강인성을 향상시키기 위해.
  • 표준 단안 3D 검출 벤치마크인 KITTI 및 nuScenes에서 최신 기술 수준의 성능을 달성하기 위해.

제안 방법

  • 카메라 시스템 파라미터(초점 거리, 수신 영역, 카메라 위치 등)를 무작위로 조작하여 기하학적 이동을 시뮬레이션하는 이미지 수준의 증강 기법을 제안한다.
  • 전경 객체를 추출하고 다른 장면에 붙여넣음으로써 새로운 훈련 인스턴스를 생성하는 인스턴스 수준의 증강 기법을 도입한다. 이는 공간 기하학을 변경하는 데 기여한다.
  • 증강 과정에서 객체 크기, 위치, 깊이 간의 기하학적 관계를 유지함으로써 물리적 타당성을 확보한다.
  • 객체 외관, 깊이, 카메라 기하학 간의 상호작용을 명시적으로 모델링하여 검출기의 일반화 능력을 향상시키는 증강 기법을 설계한다.
  • 훈련 중에 제안된 증강 기법을 적용하여, 다양한 기하학적 조건 하에서도 깊이를 일관되게 추론할 수 있는 능력을 향상시킨다.

실험 결과

연구 질문

  • RQ1카메라 파라미터의 기하학적 이동이 기존 단안 3D 검출기의 깊이 추정 신뢰성에 어떤 영향을 미치는가?
  • RQ2현재의 데이터 증강 전략이 객체 크기, 위치, 깊이 간의 기하학적 일관성을 어느 정도 유지하지 못하는가?
  • RQ3명시적인 기하학적 인지 데이터 증강이 단안 3D 검출에서 깊이 복구의 강인성을 향상시킬 수 있는가?
  • RQ4이미지 수준과 인스턴스 수준의 기하학적 증강 기법은 검출기 성능 향상에 있어 어떻게 비교되는가?
  • RQ5기하학적 관계를 유지하는 증강 기법은 KITTI 및 nuScenes 벤치마크에서 최신 기술 수준의 성능에 어떤 영향을 미치는가?

주요 결과

  • 제안된 기하학적 인지 데이터 증강 기법은 기하학적 이동 하에서 깊이 복구의 불안정성을 크게 감소시킨다.
  • KITTI 단안 3D 검출 벤치마크에서 최신 기술 수준의 성능을 달성하며, 이전 방법들을 능가한다.
  • nuScenes 벤치마크에서 3D 검출 정확도가 상당히 향상되어 일반화 능력을 확인한다.
  • 이미지 수준과 인스턴스 수준의 증강 기법을 조합하면 다양한 장면 구성에서 더 강인하고 일관된 깊이 추정이 가능해진다.
  • 증강 전략은 객체 크기, 위치, 깊이 간의 기하학적 관계를 효과적으로 유지하여 모델의 일반화 능력을 향상시킨다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.