Skip to main content
QUICK REVIEW

[논문 리뷰] Fast and Incremental Loop Closure Detection with Deep Features and Proximity Graphs

Shan An, Haogang Zhu|arXiv (Cornell University)|2020. 09. 29.
Advanced Image and Video Retrieval Techniques인용 수 6
한 줄 요약

이 논문은 단일 합성곱 신경망을 사용하여 효율적인 시각적 장소 인식을 위한 전역 및 국소 딥 특징을 추출하는 빠르고 인크리멘탈인 루프 클로징 검출 시스템인 FILD++를 제안한다. 실시간 색인을 위해 계층적 가시성 소월(HNSW) 그래프를 활용하고, 압축된 40D 국소 특징에 대해 브루트포스 매칭을 수행함으로써, 11개의 벤치마크 데이터셋 중 8개에서 최신 기준 성능을 달성하며, 대규모 뉴 콜리지 데이터셋(52,480장의 이미지)에서 평균 쿼리 시간이 22.05ms이다.

ABSTRACT

In recent years, the robotics community has extensively examined methods concerning the place recognition task within the scope of simultaneous localization and mapping applications.This article proposes an appearance-based loop closure detection pipeline named ``FILD++" (Fast and Incremental Loop closure Detection).First, the system is fed by consecutive images and, via passing them twice through a single convolutional neural network, global and local deep features are extracted.Subsequently, a hierarchical navigable small-world graph incrementally constructs a visual database representing the robot's traversed path based on the computed global features.Finally, a query image, grabbed each time step, is set to retrieve similar locations on the traversed route.An image-to-image pairing follows, which exploits local features to evaluate the spatial information. Thus, in the proposed article, we propose a single network for global and local feature extraction in contrast to our previous work (FILD), while an exhaustive search for the verification process is adopted over the generated deep local features avoiding the utilization of hash codes. Exhaustive experiments on eleven publicly available datasets exhibit the system's high performance (achieving the highest recall score on eight of them) and low execution times (22.05 ms on average in New College, which is the largest one containing 52480 images) compared to other state-of-the-art approaches.

연구 동기 및 목표

  • 시간이 오래 걸리는 특징 추출 및 해싱 기법에 대한 의존도를 줄임으로써 시각적 루프 클로징 검출의 계산 병목 현상을 해결한다.
  • 단일 딥 네트워크 내에서 전역 및 국소 특징 추출을 통합함으로써 장소 인식의 효율성과 정확도를 향상시킨다.
  • 계층적 가시성 소월(HNSW) 그래프를 이용한 인크리멘탈 색인을 통해 대규모 환경에서 실시간 성능을 달성한다.
  • 허시 코드 생성 및 복잡한 어휘 구축이 필요 없도록 압축된 딥 국소 특징에 대해 직접적인 완전 매칭을 사용함으로써 해시 기반 또는 어휘 기반 방법을 대체한다.
  • 기하학적 검증 정확도를 훼손하지 않으면서 다양한 실제 환경 데이터셋에서 높은 리콜과 낮은 실행 시간을 달성한다.

제안 방법

  • 두 번의 순방향 전파를 통해 단일 합성곱 신경망을 사용해 전역 및 국소 딥 특징을 추출함으로써 모델 복잡도와 추론 시간을 감소시킨다.
  • 전역 딥 특징을 사용해 HNSW 그래프를 인크리멘탈 방식으로 구축하여 쿼리 처리 중 빠른 유사도 검색을 가능하게 한다.
  • 전역 특징 유사도 기반으로 HNSW 그래프에서 상위-k 근접 이웃을 검색함으로써 초기 필터링을 수행한다.
  • 기하학적 검증을 위해 40차원 딥 국소 특징에 대해 브루트포스 매칭을 수행함으로써 허시 코드나 양자화를 사용하지 않는다.
  • RANSAC 기반 기하 일致성 검사를 통해 후보 이미지 쌍을 검증하여 가짜 양성(false positive)을 제거한다.
  • 전체 파이프라인을 실시간으로 작동하는 인크리멘탈 프레임워크로 통합하여 대규모 SLAM 시스템에의 구현에 적합하게 한다.
Figure 1: The modified version of DEep Local Feature (DELF) [ 61 ] architecture for feature extraction . We extract the incoming visual stream’s global and local representations via two passes of the proposed fully convolutional network. Three components constitute its structure, namely: the backbon
Figure 1: The modified version of DEep Local Feature (DELF) [ 61 ] architecture for feature extraction . We extract the incoming visual stream’s global and local representations via two passes of the proposed fully convolutional network. Three components constitute its structure, namely: the backbon

실험 결과

연구 질문

  • RQ1전역 및 국소 특징 추출을 위한 통합된 딥 네트워크가 다중 모델 또는 하이브리드 접근 방식에 비해 루프 클로징 검출의 효율성과 정확도를 향상시키는가?
  • RQ2전역 특징 색인에 HNSW 그래프를 사용할 경우 대규모 시각 기반 데이터베이스에서 검색 속도와 리콜에 어떤 영향을 미치는가?
  • RQ3압축된 딥 국소 특징에 대해 브루트포스 매칭이 속도와 정확도 측면에서 해시 기반 또는 어휘 기반 방법을 얼마나 효과적으로 대체할 수 있는가?
  • RQ4허시 코드 생성 단계를 제거함으로써 실행 시간과 시스템 단순성 측면에서 어떤 성능 향상이 이루어지는가?
  • RQ5다양한 실제 환경 데이터셋에서 FILD++는 최신 기준 방법에 비해 리콜과 지연 시간 측면에서 어떻게 비교되는가?

주요 결과

  • FILD++는 공개된 11개 데이터셋 중 8개에서 가장 높은 리콜 점수를 기록하여 뛰어난 검출 정확도를 입증한다.
  • 뉴 콜리지 데이터셋(52,480장의 이미지)에서 쿼리당 평균 처리 시간이 22.05ms로, 이전의 FILD 방법(50.28ms/쿼리)보다 현저히 빠르게 된다.
  • 뉴 콜리지 데이터셋에서 특징 추출 시간은 17.69ms에서 14.62ms로 단축되었으며, 허시 코드 생성 단계가 필요 없었다.
  • RANSAC 검증 단계의 처리 시간은 7.55ms에서 1.72ms로 단축되어 전체 성능 향상에 기여했다.
  • 40차원 딥 국소 특징의 사용으로 CasHash가 필요 없이 효율적인 브루트포스 매칭이 가능해져 파이프라인을 단순화하고 계산 오버헤드를 감소시켰다.
  • 딥 특징의 분류 능력과 효과적인 기하학적 검증 덕분에 다양한 환경, 특히 고도로 무늬가 있는 장면이나 동적인 장면에서도 높은 정확도와 강건성을 유지한다.
Figure 2: An overview of the proposed loop closure detection pipeline. Global and local Convolution Neural Network (CNN) -based features are extracted as the incoming image stream enters the system. The global features enter the First-In-First-Out (FIFO) queue, and subsequently, they are fed into th
Figure 2: An overview of the proposed loop closure detection pipeline. Global and local Convolution Neural Network (CNN) -based features are extracted as the incoming image stream enters the system. The global features enter the First-In-First-Out (FIFO) queue, and subsequently, they are fed into th

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.