Skip to main content
QUICK REVIEW

[논문 리뷰] AA-RMVSNet: Adaptive Aggregation Recurrent Multi-view Stereo Network

Zizhuang Wei, Qingtian Zhu|arXiv (Cornell University)|2021. 08. 09.
Advanced Vision and Imaging참고 문헌 33인용 수 13
한 줄 요약

AA-RMVSNet는 적응형 집합 모듈을 갖춘 반복적 다중 시야 입체망을 제안한다: 시야 내 모듈은 맥락 인지 다중 척도 특징 추출을 위해, 시야 간 모듈은 막힘 처리를 위해 픽셀 단위의 시야 선택을 위해 사용된다. 이는 최신 기술 수준의 성능을 달성하여 탱크스 앤 템플스 벤치마크에서 61.51의 평균 F-스코어로 1위를 기록하고 DTU에서도 경쟁력 있는 결과를 보이며 3차원 재구성의 높은 정확도와 완전성을 입증한다.

ABSTRACT

In this paper, we present a novel recurrent multi-view stereo network based on long short-term memory (LSTM) with adaptive aggregation, namely AA-RMVSNet. We firstly introduce an intra-view aggregation module to adaptively extract image features by using context-aware convolution and multi-scale aggregation, which efficiently improves the performance on challenging regions, such as thin objects and large low-textured surfaces. To overcome the difficulty of varying occlusion in complex scenes, we propose an inter-view cost volume aggregation module for adaptive pixel-wise view aggregation, which is able to preserve better-matched pairs among all views. The two proposed adaptive aggregation modules are lightweight, effective and complementary regarding improving the accuracy and completeness of 3D reconstruction. Instead of conventional 3D CNNs, we utilize a hybrid network with recurrent structure for cost volume regularization, which allows high-resolution reconstruction and finer hypothetical plane sweep. The proposed network is trained end-to-end and achieves excellent performance on various datasets. It ranks $1^{st}$ among all submissions on Tanks and Temples benchmark and achieves competitive results on DTU dataset, which exhibits strong generalizability and robustness. Implementation of our method is available at https://github.com/QT-Zhu/AA-RMVSNet.

연구 동기 및 목표

  • 딥 러닝을 사용한 얇은 구조물과 낮은 무늬 표면의 3차원 재구성 도전 과제를 해결하기 위해.
  • 다양한 막힘 상황에서 적응형 시야 집합을 통해 복잡한 환경에서의 강인성을 향상시키기 위해.
  • 하이브리드 LSTM 아키텍처를 통해 고해상도 재구성을 유지하면서 메모리 소비를 줄이기 위해.
  • 깊이 맵 정확도와 완전성을 향상시키기 위해 특징 표현 및 비용 볼륨 정규화를 개선하기 위해.

제안 방법

  • 왜곡 가능한 컨볼루션과 다중 척도 특징 융합을 사용한 시야 내 적응형 집합 모듈을 도입하여 맥락 인지 특징 추출을 향상시킨다.
  • 픽셀 단위의 주의 맵을 생성하여 잘 매칭된 시야 쌍을 우선시하는 시야 간 비용 볼륨 집합 모듈을 제안한다.
  • 비용 볼륨 정규화를 위해 반복적 LSTM 아키텍처를 가진 하이브리드 네트워크를 활용하여 고해상도 및 세밀한 평면 스위프트를 가능하게 한다.
  • 경량 모듈을 사용한 엔드 투 엔드 학습을 통해 성능 향상과 함께 메모리 효율성을 유지한다.
  • 특징 추출 및 시야 선택 단계에서 모두 적응형 집합을 적용하여 다양한 시나리오 유형에서 강인성을 향상시킨다.
Figure 1: Illustration of multi-view 3D reconstruction of Scan 77 in DTU dataset [ 3 ] using the proposed AA-RMVSNet. (a) The reference image; (b) adaptive sampling locations in our intra-view AA approach; (c) the depth map estimated by AA-RMVSNet after filtering; (d) the recovered dense 3D model.
Figure 1: Illustration of multi-view 3D reconstruction of Scan 77 in DTU dataset [ 3 ] using the proposed AA-RMVSNet. (a) The reference image; (b) adaptive sampling locations in our intra-view AA approach; (c) the depth map estimated by AA-RMVSNet after filtering; (d) the recovered dense 3D model.

실험 결과

연구 질문

  • RQ1적응형 시야 내 특징 집합이 얇고 무문자 구조물의 깊이 추정에 개선을 이룰 수 있는가?
  • RQ2주목적 맵을 통한 픽셀 단위의 시야 간 시야 선택이 복잡한 환경에서 막힘의 영향을 줄일 수 있는가?
  • RQ3반복적 LSTM 기반 비용 볼륨 정규화가 정확도와 메모리 효율성 면에서 3D CNN보다 뛰어나게 작용하는가?
  • RQ4결합되었을 때 제안된 적응형 모듈의 성능 및 메모리 비용은 어떻게 비교되는가?

주요 결과

  • AA-RMVSNet는 탱크스 앤 템플스 벤치마크에서 61.51의 평균 F-스코어를 기록하여 모든 제출물 중 1위를 차지했다.
  • DTU 데이터셋에서 전체 평균 깊이 오차가 0.357 mm로 기준선 대비 8.7% 향상되었다.
  • 시야 내 적응형 집합은 완전성을 0.28 향상시키면서 오직 1.74 GB의 메모리만 추가로 소비했고, 시야 간 AA는 정확도를 0.31 향상시키며 0.11 GB의 오버헤드를 유발했다.
  • 전체 모델는 오직 4.25 GB의 메모리만 사용하여 800×600 해상도의 고해상도 깊이 맵을 생성하여 강력한 메모리 효율성을 입증했다.
  • 시각화 결과는 R-MVSNet 및 기타 기준선 대비 얇은 물체와 막힘 영역 재구성에서 뚜렷한 향상을 보였다.
Figure 2: Overall architecture of AA-RMVSNet that consists of 4 stages. Intra-view AA module aims to aggregate context-aware features for multiple scales and regions with varying richness of texture. Inter-view AA module adaptively aggregates cost volumes of different views by yielding pixel-wise at
Figure 2: Overall architecture of AA-RMVSNet that consists of 4 stages. Intra-view AA module aims to aggregate context-aware features for multiple scales and regions with varying richness of texture. Inter-view AA module adaptively aggregates cost volumes of different views by yielding pixel-wise at

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.