[논문 리뷰] Deep AutoEncoder-based Lossy Geometry Compression for Point Clouds
점구름의 손실형 지오메트리 압축을 위한 딥 자동 인코더 아키텍처를 제안하여 점구름을 직접 입력으로 사용하고 MPEG TMC13 대비 우수한 rate–distortion 성능을 달성합니다(평균 73.15% BD-rate 혜택).
Point cloud is a fundamental 3D representation which is widely used in real world applications such as autonomous driving. As a newly-developed media format which is characterized by complexity and irregularity, point cloud creates a need for compression algorithms which are more flexible than existing codecs. Recently, autoencoders(AEs) have shown their effectiveness in many visual analysis tasks as well as image compression, which inspires us to employ it in point cloud compression. In this paper, we propose a general autoencoder-based architecture for lossy geometry point cloud compression. To the best of our knowledge, it is the first autoencoder-based geometry compression codec that directly takes point clouds as input rather than voxel grids or collections of images. Compared with handcrafted codecs, this approach adapts much more quickly to previously unseen media contents and media formats, meanwhile achieving competitive performance. Our architecture consists of a pointnet-based encoder, a uniform quantizer, an entropy estimation block and a nonlinear synthesis transformation module. In lossy geometry compression of point cloud, results show that the proposed method outperforms the test model for categories 1 and 3 (TMC13) published by MPEG-3DG group on the 125th meeting, and on average a 73.15\% BD-rate gain is achieved.
연구 동기 및 목표
- 불규칙한 3D 점군에 대한 유연하고 데이터 기반의 지오메트리 압축을 고무하고 가능하게 한다.
- 보셀화나 이미지 유사 표현 없이 원시 점군을 직접 처리하는 엔드-투-엔드 autoencoder 프레임워크를 개발한다.
- 학습 가능한 엔트로피 모델과 미분 가능 양자화 스킴을 도입하여 rate–distortion를 최적화한다.
- 일반 객체 카테고리 전반에서 MPEG PCC 베이스라인에 대해 경쟁력 있는 압축 성능을 입증한다.
제안 방법
- 순서를 불변인 잠재 표현을 얻기 위해 PointNet 기반 인코더를 사용한다.
- 직선 전파(straight-through) 또는 균일 잡음 근사를 학습 중에 사용하여 이산 잠재 코드를 생성하는 균일 양자화를 적용한다.
- 비트레이트 추정용 잠재 코드의 사전 분포를 모델링하기 위해 엔트로피 추정(엔트로피 보틀넥) 모듈을 도입한다.
- Chamfer 거리 왜곡과 추정 비트레이트를 결합한 rate–distortion 목표를 최적화한다.
- 잠재 코드로부터 3D 점군을 재구성하기 위해 완전 연결된 디코더를 사용한다.
- ShapeNet에서 학습하고 평가하여 다수의 객체 카테고리에서 MPEG TMC13과 비교한다.
실험 결과
연구 질문
- RQ1엔드-투-엔드 autoencoder가 보셀화 없이 점군으로부터 직접 효율적인 손실 있는 지오메트리 표현을 학습할 수 있는가?
- RQ2학습된 엔트로피 모델과 미분 가능 양자화를 사용할 때 MPEG PCC의 TMC13 대비 어느 정도의 rate–distortion 성능 향상을 얻을 수 있는가?
- RQ3엔트로피 보틀넥의 도입이 점군 지오메트리의 비트레이트와 재구성 품질에 어떤 영향을 미치는가?
- RQ4다양한 비트레이트에서 chair, airplane, table, car 등 서로 다른 객체 카테고리에서 이 접근법이 견고한가?
주요 결과
- 제안된 autoencoder 기반 지오메트리 코덱은 chair, airplane, table, car 카테고리에서 모든 테스트 비트레이트에서 TMC13을 능가한다.
- 평균적으로 본 방법은 TMC13 대비 73.15% BD-rate 이득을 달성한다.
- 아블레이션 연구는 엔트로피 보틀넥 모듈이 엔트로피 추정이 없는 베이스라인 대비 19.3% BD-rate 이득을 제공한다는 것을 보여준다.
- 재구성된 점군은 밀도가 더 촘촘하고 PSNR이 비슷할 때 포인트당 비트 수가 더 낮다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.