Skip to main content
QUICK REVIEW

[논문 리뷰] ACORN: Adaptive Coordinate Networks for Neural Scene Representation

Julien Martel, David B. Lindell|arXiv (Cornell University)|2021. 05. 06.
3D Shape Modeling and Analysis참고 문헌 4인용 수 12
한 줄 요약

ACORN은 2D 및 3D 시각을 다중스케일 블록으로 적응적으로 분할하는 하이브리드 암시-명시 신경망 아키텍처를 도입하여 좌표 인코더를 사용해 특징 격자를 생성하고 경량 디코더를 통해 효율적인 추론을 구현한다. 이는 약 40 dB PSNR에서 최초로 기가픽셀 이미지 피팅을 달성하였으며, 3D 형태 학습 시간을 일수에서 분으로 단축시키고 메모리 사용량을 10배 이상 감소시켰다.

ABSTRACT

Neural representations have emerged as a new paradigm for applications in rendering, imaging, geometric modeling, and simulation. Compared to traditional representations such as meshes, point clouds, or volumes they can be flexibly incorporated into differentiable learning-based pipelines. While recent improvements to neural representations now make it possible to represent signals with fine details at moderate resolutions (e.g., for images and 3D shapes), adequately representing large-scale or complex scenes has proven a challenge. Current neural representations fail to accurately represent images at resolutions greater than a megapixel or 3D scenes with more than a few hundred thousand polygons. Here, we introduce a new hybrid implicit-explicit network architecture and training strategy that adaptively allocates resources during training and inference based on the local complexity of a signal of interest. Our approach uses a multiscale block-coordinate decomposition, similar to a quadtree or octree, that is optimized during training. The network architecture operates in two stages: using the bulk of the network parameters, a coordinate encoder generates a feature grid in a single forward pass. Then, hundreds or thousands of samples within each block can be efficiently evaluated using a lightweight feature decoder. With this hybrid implicit-explicit network architecture, we demonstrate the first experiments that fit gigapixel images to nearly 40 dB peak signal-to-noise ratio. Notably this represents an increase in scale of over 1000x compared to the resolution of previously demonstrated image-fitting experiments. Moreover, our approach is able to represent 3D shapes significantly faster and better than previous techniques; it reduces training times from days to hours or minutes and memory requirements by over an order of magnitude.

연구 동기 및 목표

  • 큰 스케일 또는 복잡한 3D 시각 및 고해상도 이미지를 처리하는 데 있어 기존 신경 시각 표현의 확장성 한계를 해결하기 위해.
  • 하이브리드 암시-명시 아키텍처를 도입하여 신경 표현에서 계산 효율성과 메모리 사용 간의 상충 관계를 극복하기 위해.
  • 지역 신호 복잡도에 따라 학습 및 추론 중 자원 할당을 적응적으로 조정하여 정확도와 효율성을 동시에 향상시키기 위해.
  • 신경망을 사용하여 고해상도 이미지 및 복잡한 3D 기하 구조의 피팅에서 최고 성능을 달성하기 위해.

제안 방법

  • 메모리 구조가 쿼드트리나 옥트리와 유사한 다중스케일 블록-좌표 분해를 사용하며, 이는 학습 도중 신호 영역을 적응적으로 하위 분할하기 위해 최적화된다.
  • 좌표 인코더는 각 블록에 대해 단일 순방향 전파를 통해 저해상도 특징 격자를 생성하며, 네트워크 파ameter의 대부분을 사용한다.
  • 경량 특징 디코더를 통해 각 블록 내 연속적 좌표의 평가를 보다 빠르고 효율적으로 수행할 수 있다.
  • 블록 분할과 네트워크 가중치를 동시에 최적화하기 위해 정수선형계획을 활용한 고유한 학습 루틴을 사용한다.
  • 특징 격자와 연속적 디코딩을 조합함으로써 암시적(좌표 기반) 및 명시적(격자 기반) 표현을 모두 지원한다.
  • 학습 도중 자원 할당을 동적으로 조정하여 지역 복잡도가 높은 영역에 더 많은 블록을 할당한다.

실험 결과

연구 질문

  • RQ1신경 시각 표현은 고해상도 기가픽셀 이미지를 유지하면서도 효율적인 추론이 가능한가?
  • RQ2복잡한 3D 시각 모델링에서 신경 표현이 메모리 효율성과 계산 속도를 어떻게 균형 잡을 수 있는가?
  • RQ3고정 격자 또는 균일 해상도 방법에 비해 적응형 블록 분할이 표현 정확도와 학습 효율성을 향상시킬 수 있는가?
  • RQ4하이브리드 암시-명시 아키텍처는 순수 암시적 또는 명시적 신경 표현에 비해 학습 시간, 메모리 사용량, 재구성 품질 측면에서 얼마나 뛰어나게 성능을 발휘할 수 있는가?

주요 결과

  • ACORN은 8192×8192 해상도의 기가픽셀 이미지(약 40 dB PSNR)를 신경 표현에 성공적으로 피팅하는 데에 최초로 성공하였다.
  • 3D 형태 표현의 학습 시간을 일수에서 분으로 단축시켜 이전 방법 대비 메모리 사용량을 10배 이상 감소시켰다.
  • Lucy(1400만 개 정점)와 같은 3D 모델의 경우 ACORN은 단지 68MB의 모델 크기로 압축되지 않은 1.2GB 및 압축된 380MB 버전보다 훨씬 작았다.
  • 경계 영역에서도 높은 정밀도의 피팅 성능를 보였으며, 연속성 강제 조건 없이 일부 경우에만 미세한 아티팩트가 관찰되었다.
  • 기존의 이미지 피팅 실험은 메가픽셀 이하 해상도로 제한되었으나, ACORN은 이에 비해 해상도 스케일을 1000배 향상시켰다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.