[논문 리뷰] LookinGood: Enhancing Performance Capture with Real-time Neural Re-Rendering
LookinGood는 실시간 신경 재렌더링 시스템을 소개하며, 성능 캡처 파이프라인에서 생성된 저품질 출력물을 복합적으로 보완함으로써 품질을 햖थ고, 깊이 학습 모델을 사용해 동시에 완성, 초해상도 처리, 노이즈 제거를 수행한다. 의미적 주목도와 이방성 제거에 중점을 둔 자기지도 학습 손실을 통해 훈련된 이 방법은 VR/AR 응용 분야에서 고해상도, 시간적으로 안정적이고 스테레오 일관성이 확보된 렌더링을 달성하며, 사용자 연구에서 원본 캡처 결과를 능가한다. 또한, 훈련 데이터에 포함되지 않은 주제들에도 일반화 가능하다.
Motivated by augmented and virtual reality applications such as telepresence, there has been a recent focus in real-time performance capture of humans under motion. However, given the real-time constraint, these systems often suffer from artifacts in geometry and texture such as holes and noise in the final rendering, poor lighting, and low-resolution textures. We take the novel approach to augment such real-time performance capture systems with a deep architecture that takes a rendering from an arbitrary viewpoint, and jointly performs completion, super resolution, and denoising of the imagery in real-time. We call this approach neural (re-)rendering, and our live system "LookinGood". Our deep architecture is trained to produce high resolution and high quality images from a coarse rendering in real-time. First, we propose a self-supervised training method that does not require manual ground-truth annotation. We contribute a specialized reconstruction error that uses semantic information to focus on relevant parts of the subject, e.g. the face. We also introduce a salient reweighing scheme of the loss function that is able to discard outliers. We specifically design the system for virtual and augmented reality headsets where the consistency between the left and right eye plays a crucial role in the final user experience. Finally, we generate temporally stable results by explicitly minimizing the difference between two consecutive frames. We tested the proposed system in two different scenarios: one involving a single RGB-D sensor, and upper body reconstruction of an actor, the second consisting of full body 360 degree capture. Through extensive experimentation, we demonstrate how our system generalizes across unseen sequences and subjects. The supplementary video is available at http://youtu.be/Md3tdAKoLGU.
연구 동기 및 목표
- AR/VR 응용 분야에서 발생하는 시각적 결함(공백, 노이즈, 해상도 저하, 낮은 조명 등)을 해결하기 위해.
- 수동 애너테이션 없이 3D 재구성에서 유도된 거친 2D 렌더링을 향상시키는 실시간 신경 재렌더링 시스템을 개발하기 위해.
- 특히 이중 시력 렌더링에서 최적의 VR/AR 사용자 경험을 확보하기 위해 시간적 안정성과 스테레오 일관성을 보장하기 위해.
- 의미 인식 손실 함수를 활용한 자기지도 학습을 통해 훈련 데이터에 포함되지 않은 주제와 시퀀스에 일반화할 수 있도록 하기 위해.
제안 방법
- 저해상도이자 결함이 있는 2D 렌더링을 실시간으로 고품질, 고해상도 출력으로 매핑하는 딥 네트워크를 훈련한다.
- 지상 진실 애너테이션의 필요성을 제거하기 위해 재구성 손실에 의미적 감시를 통합한 자기지도 학습 접근법을 사용한다.
- 손실 함수 내에서 주목도 기반 재가중 전략을 통해 얼굴과 같은 핵심 영역을 강조하고 외곽선을 억제한다.
- 모델 양자화(16비트 부동소수점)와 효율적인 아키텍처 설계를 통해 실시간 추론을 최적화하여 NVIDIA Titan V에서 스테레오 쌍당 29ms의 성능을 달성한다.
- 예측된 재렌더링 결과의 프레임 간 차이를 최소화하여 시간적 안정성을 확보한다.
- 테스트 시점에 주제를 배경에서 분리하는 이진 세그먼테이션 마스크도 예측한다.
실험 결과
연구 질문
- RQ1수동 애너테이션 없이 자기지도 학습 방식으로 저품질 실시간 성능 캡처 출력을 효과적으로 향상시킬 수 있는가?
- RQ2손실 함수에 의미 정보를 어떻게 통합하여 얼굴과 같은 핵심 영역의 시각적 품질을 향상시킬 수 있는가?
- RQ3신경 재렌더링 시스템이 VR/AR 응용 분야에서 시간적 안정성과 스테레오 일관성을 어느 정도 확보할 수 있는가?
- RQ4학습 데이터에 포함되지 않은 주제와 시퀀스에 대해 시스템의 일반화 능력은 얼마나 뛰어나며, 성능 저하 정도는 어떠한가?
주요 결과
- 단일 GPU에서 스테레오 쌍당 29ms의 실시간 성능을 달성하여 VR/AR 응용 분야의 요구 조건을 충족한다.
- 사용자 연구에서 참가자 65%가 신경 재렌더링 출력을 지상 진실보다 선호했으며, 해상도는 약간 낮지만 더 안정적인 마스크를 제공한다고 평가했다.
- 학습 데이터에 포함되지 않은 주제와 시퀀스에 대해서도 잘 일반화되며, 심각하게 손상된 입력에서는 유일하게 점진적인 품질 저하가 발생한다.
- 주목도 재가중을 통한 자기지도 손실은 얼굴 품질 향상과 외곽선 영향 감소에 크게 기여한다.
- 재렌더링 출력의 프레임 간 변동을 최소화함으로써 시간적으로 안정된 결과를 생성한다.
- 입력이 심각하게 손상된 경우 모델은 흐릿한 결과를 환영하는 경향이 있어 극단적인 경우의 주요 한계를 보여준다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.