Skip to main content
QUICK REVIEW

[논문 리뷰] RoMa: Robust Dense Feature Matching

Johan Edstedt, Qiyu Sun|arXiv (Cornell University)|2023. 05. 24.
Human Pose and Action Recognition인용 수 5
한 줄 요약

RoMa는 극한의 실생활 변화에서 강건한 밀도 높은 특징 매칭을 위해 고정된 DINOv2 특징을 조잡한 수준의 매칭에 사용하고, 전문화된 ConvNet 특징을 세밀한 수준의 보정에 사용하는 강력한 밀도 높은 특징 매칭 방법을 제안한다. 변환기 디코더를 통해 잠재적 기준점 확률을 예측하여 다중모달 분포를 모델링한다. 이는 도전적인 WxBS 벤치마크에서 36% 향상된 성능을 기록하며, 극한의 실생활 변화에 대한 밀도 높은 매칭의 새로운 최고 성능을 수립한다.

ABSTRACT

Feature matching is an important computer vision task that involves estimating correspondences between two images of a 3D scene, and dense methods estimate all such correspondences. The aim is to learn a robust model, i.e., a model able to match under challenging real-world changes. In this work, we propose such a model, leveraging frozen pretrained features from the foundation model DINOv2. Although these features are significantly more robust than local features trained from scratch, they are inherently coarse. We therefore combine them with specialized ConvNet fine features, creating a precisely localizable feature pyramid. To further improve robustness, we propose a tailored transformer match decoder that predicts anchor probabilities, which enables it to express multimodality. Finally, we propose an improved loss formulation through regression-by-classification with subsequent robust regression. We conduct a comprehensive set of experiments that show that our method, RoMa, achieves significant gains, setting a new state-of-the-art. In particular, we achieve a 36% improvement on the extremely challenging WxBS benchmark. Code is provided at https://github.com/Parskatt/RoMa

연구 동기 및 목표

  • 조명, 시점, 스케일, 질감 변화와 같은 극한의 실생활 변화에서 밀도 높은 특징 매칭의 부족한 강건성 해결
  • DINOv2에서 사전 학습된 고정된 자기지도 학습 특징을 활용하여 미세조정 또는 랜덤 초기화된 특징의 한계를 극복하고 일반화 능력 향상
  • 기초 모델의 조잡한 DINOv2 특징과 세밀한 국소화에 적합한 전문화된 ConvNet 특징을 조합하여 매칭 정확도 향상
  • 기준점 확률 예측을 통해 좌표 회귀 대신 매칭을 모델링하여 다중모달 매칭 시나리오에서의 강건성 향상
  • 이중 손실 전략—조잡한 매칭에 대한 분류 기반 회귀와 보정에 대한 강건한 회귀—을 통해 데이터 분포와 더 잘 일치하는 훈련 최적화

제안 방법

  • 제한된 3D 지도 데이터에서 오버피팅을 방지하기 위해 고정된 DINOv2 특징을 조잡한 특징 추출을 위한 강력하고 일반적인 백본으로 사용
  • 정확한 국소화를 가능하게 하는 세밀한 특징을 추출하기 위해 DINOv2 특징과 별도로 훈련된 전문화된 ConvNet 도입
  • DINOv2의 조잡한 특징과 ConvNet의 세밀한 특징을 융합하여 특징 피라미드를 구성함으로써 다중 해상도, 정확한 매칭 가능
  • 좌표에 의존하지 않는 변환기 디코더 설계로 기준점 확률을 예측하여 다중모달 매칭 분포 표현 가능
  • 이중 단계 손실 전략 구현: 다중모달 분포를 다루기 위해 조잡한 매칭에 대한 분류 기반 회귀, 이후 단일모달 국소 분포를 다루기 위한 강건한 회귀
  • 전체 파이프라인의 엔드 투 엔드 최적화를 가능하게 하기 위해 지도 기반 대응 데이터와 제안된 손실 함수의 조합을 사용해 모델 훈련
Figure 2 : Illustration of our robust approach RoMa. Our contributions are shown with green highlighting and a checkmark, while previous approaches are indicated with gray highlights and a cross. Our first contribution is using a frozen foundation model for coarse features, compared to fine-tuning o
Figure 2 : Illustration of our robust approach RoMa. Our contributions are shown with green highlighting and a checkmark, while previous approaches are indicated with gray highlights and a cross. Our first contribution is using a frozen foundation model for coarse features, compared to fine-tuning o

실험 결과

연구 질문

  • RQ1DINOv2에서 유도된 고정된 자기지도 학습 특징은 미세조정 또는 무작위 초기화된 특징에 비해 밀도 높은 특징 매칭의 강건성에 상당한 기여를 할 수 있는가?
  • RQ2기초 모델의 조잡한 특징과 전문화된 ConvNet의 세밀한 특징을 융합하는 것이 특징을 함께 훈련하는 것보다 더 나은 국소화 정확도를 제공하는가?
  • RQ3직접 좌표 회귀 대신 기준점 확률 예측을 통해 매칭을 모델링하는 것이 다중모달 매칭 시나리오에서 성능 향상에 기여하는가?
  • RQ4이중 손실 전략—조잡한 매칭에 대한 분류 기반 회귀와 보정에 대한 강건한 회귀—은 표준 L2 손실에 비해 더 효과적인가?
  • RQ5제안된 방법은 이전 방법이 실패하는 극한의 벤치마크, 예를 들어 WxBS에서 어느 정도 성능 향상을 이룰 수 있는가?

주요 결과

  • RoMa는 극도로 도전적인 WxBS 벤치마크에서 평균 정밀도(mAP) 기준 36% 상대적 향상을 기록하며 새로운 최고 성능 달성
  • IMC2022 벤치마크에서 이전 최고 성능 방법 대비 상대 오차를 26% 감소시켜 넓은 기준선 매칭에서 강력한 일반화 능력 입증
  • ScanNet-1500 벤치마크에서 AUC@20° 점수가 70 이상을 처음으로 기록하여 저질감, 고변동 내부 환경에서 뛰어난 성능 입증
  • MegaDepth-8-Scenes 벤치마크에서 AUC@20°는 85.3을 기록하며 이전 최고 성능 방법인 DKM(84.2)과 ASpanFormer(82.9)을 초월
  • InLoc 벤치마크에서 시각적 국소화 작업에서 0.25m 이내 및 2° 이내 정확도 89.9%를 달성하여 이전 모든 방법을 뛰어넘음
  • 더 복잡한 아키텍처를 도입했음에도 불구하고 추론 시간은 7%만 증가(560×560 해상도 기준 198.8ms 대비 186.3ms)하여 효율적인 구현 가능
Figure 3 : Illustration of localizability of matches. At infinite resolution the match distribution can be seen as a 2D surface (illustrated as 1D lines in the figure), however at a coarser scale $s$ this distribution becomes blurred due to motion boundaries. This means it is necessary to both use a
Figure 3 : Illustration of localizability of matches. At infinite resolution the match distribution can be seen as a 2D surface (illustrated as 1D lines in the figure), however at a coarser scale $s$ this distribution becomes blurred due to motion boundaries. This means it is necessary to both use a

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.