Skip to main content
QUICK REVIEW

[논문 리뷰] Distilled Low Rank Neural Radiance Field with Quantization for Light Field Compression

Jinglei Shi, Christine Guillemot|arXiv (Cornell University)|2022. 07. 30.
Advanced Vision and Imaging인용 수 5
한 줄 요약

이 논문은 빛의 장을 압축하기 위한 양자화된 휘발성 저질서 신경 렌더링 필드(QDLR-NeRF)를 제안한다. ADMM를 통한 저질서 제약 적용, 텐서 트리플렛 분해를 통한 모델 휘발성, 그리고 전역 코드북을 사용한 점진적 가중치 양자화를 통해 원본 NeRF 모델 크기를 3.3%로 감소시키면서도 높은 시각 합성 품질을 유지하며, 비율-왜곡 효율성에서 최신 압축 기법들을 능가한다.

ABSTRACT

We propose in this paper a Quantized Distilled Low-Rank Neural Radiance Field (QDLR-NeRF) representation for the task of light field compression. While existing compression methods encode the set of light field sub-aperture images, our proposed method learns an implicit scene representation in the form of a Neural Radiance Field (NeRF), which also enables view synthesis. To reduce its size, the model is first learned under a Low-Rank (LR) constraint using a Tensor Train (TT) decomposition within an Alternating Direction Method of Multipliers (ADMM) optimization framework. To further reduce the model's size, the components of the tensor train decomposition need to be quantized. However, simultaneously considering the optimization of the NeRF model with both the low-rank constraint and rate-constrained weight quantization is challenging. To address this difficulty, we introduce a network distillation operation that separates the low-rank approximation and the weight quantization during network training. The information from the initial LR-constrained NeRF (LR-NeRF) is distilled into a model of much smaller dimension (DLR-NeRF) based on the TT decomposition of the LR-NeRF. We then learn an optimized global codebook to quantize all TT components, producing the final QDLR-NeRF. Experimental results show that our proposed method yields better compression efficiency compared to state-of-the-art methods, and it additionally has the advantage of allowing the synthesis of any light field view with high quality.

연구 동기 및 목표

  • 빛의 장 촬영에서 높은 데이터 볼륨 문제를 해결하기 위해 압축을 위한 컴act하고 고정밀한 표현 방식을 개발한다.
  • NeRF 기반 빛의 장 압축에서 저질서 근사와 비율 제약이 있는 가중치 양자화를 동시에 최적화하는 데 어려움을 해결한다.
  • 개별 뷰가 아닌 NeRF 모델 자체를 압축하여 빛의 장 데이터의 효율적 전송을 가능하게 한다.
  • 지식 휘발성과 점진적 양자화를 통해 극단적인 모델 압축에도 불구하고 높은 시각 합성 품질을 유지한다.

제안 방법

  • 먼저, 저질서 제약 하에 텐서 랭크를 최적화하기 위해 분할 최적화 방법(ADMM)을 사용하여 저질서 제약이 가해진 NeRF(LR-NeRF)를 훈련한다.
  • LR-NeRF의 지식을 텐서 트리플렛(TT) 분해를 통해 더 작은, 휘발성 저질서 NeRF(DLR-NeRF)로 이관하여 파라미터 수를 감소시킨다.
  • 모든 레이어의 TT 구성 요소를 점진적으로 양자화하기 위해 k-means 클러스터링을 통해 전역 코드북을 학습한다. 이 과정은 첫 번째 레이어부터 마지막 레이어까지 순차적으로 진행된다.
  • 공유된 코드북을 사용하여 레이어별로 양자화 과정을 적용함으로써 재구성 오차를 최소화하면서 모델 크기를 극적으로 감소시킨다.
  • 전체 파이프라인은 휘발성과 저질서 최적화를 분리함으로써 안정적인 훈련과 효과적인 압축을 가능하게 한다.

실험 결과

연구 질문

  • RQ1NeRF 기반 빛의 장 압축 프레임워크에서 저질서 근사와 가중치 양자화를 효과적으로 융합할 수 있는가?
  • RQ2저질서 제약이 가해진 NeRF에서 유의미한 품질 손실 없이 매우 작은 압축 모델로 지식을 전달할 수 있는가?
  • RQ3전역 코드북을 사용한 점진적 양자화가 모델 크기와 시각 합성 품질에 미치는 영향는 어떠한가?
  • RQ4비율-왜곡 성능 측면에서 제안된 방법은 최신 빛의 장 압축 기법들과 비교해 어떻게 성능을 냈는가?

주요 결과

  • 저질서 제약, 휘발성, 양자화를 거친 후 모델 크기가 원본 NeRF의 100%에서 3.3%로 감소하여 30배의 크기 감소를 달성했다.
  • 원본 NeRF에서 LR-NeRF로의 전환 시 PSNR가 약 3dB 감소하였는데, 주로 추정된 카메라 포즈와 저질서 근사에 기인한다.
  • LR-NeRF를 DLR-NeRF로 휘발성함으로써 모델 크기를 3배(16.7%)로 감소시켰고, PSNR 손실은 0.1dB에 불과하여 저질서 지식 전달이 효과적임을 보여준다.
  • 점진적 양자화를 통해 모델 크기를 원본의 3.3%로 줄였고, PSNR 손실은 0.6dB에 그쳐 전역 코드북의 효과를 입증한다.
  • 합성 빛의 장에서는 기준 기법들보다 유의미하게 뛰어난 비율-왜곡 성능을 달성하였으며, 실세계의 Lytro 촬영 데이터보다는 깨끗한 데이터에서 더 큰 성능 향상을 보였다.
  • NeRF의 암묵적 장면 표현 덕분에 고밀도 빛의 장 뷰라도 고품질로 합성 가능하다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.