[논문 리뷰] Deep Learning-Based Video Coding: A Review and A Case Study
이 논문은 이미지/비디오 코딩을 위한 딥러닝 접근법을 검토하고 이를 딥 스킴(Deep schemes)과 딥 툴(Deep tools)로 분류하며 CNN-based in-loop filtering과 CNN-BARC가 HEVC 대비 비트레이트 절감 효과를 크게 보이는 DLVC 프로토타입을 제시한다.
The past decade has witnessed great success of deep learning technology in many disciplines, especially in computer vision and image processing. However, deep learning-based video coding remains in its infancy. This paper reviews the representative works about using deep learning for image/video coding, which has been an actively developing research area since the year of 2015. We divide the related works into two categories: new coding schemes that are built primarily upon deep networks (deep schemes), and deep network-based coding tools (deep tools) that shall be used within traditional coding schemes or together with traditional coding tools. For deep schemes, pixel probability modeling and auto-encoder are the two approaches, that can be viewed as predictive coding scheme and transform coding scheme, respectively. For deep tools, there have been several proposed techniques using deep learning to perform intra-picture prediction, inter-picture prediction, cross-channel prediction, probability distribution prediction, transform, post- or in-loop filtering, down- and up-sampling, as well as encoding optimizations. In the hope of advocating the research of deep learning-based video coding, we present a case study of our developed prototype video codec, namely Deep Learning Video Coding (DLVC). DLVC features two deep tools that are both based on convolutional neural network (CNN), namely CNN-based in-loop filter (CNN-ILF) and CNN-based block adaptive resolution coding (CNN-BARC). Both tools help improve the compression efficiency by a significant margin. With the two deep tools as well as other non-deep coding tools, DLVC is able to achieve on average 39.6\% and 33.0\% bits saving than HEVC, under random-access and low-delay configurations, respectively. The source code of DLVC has been released for future researches.
연구 동기 및 목표
- 이미지/비디오 코딩 분야의 딥러닝 현황을 2018년까지 조사한다.
- 딥 스킴과 딥 툴의 구분과 장단점을 분석한다.
- 이득을 설명하기 위해 실용적인 Deep Learning Video Coding (DLVC) 프로토타입을 시연한다.
제안 방법
- 픽셀 확률 모델링과 오토인코더 접근법을 이용한 대표적 딥 코딩 연구를 검토한다.
- 딥 스킴을 예측/픽셀 확률 모델 또는 변환 기반 오토인코더로 분류하고 전통 코덱 내에서 사용되는 딥 툴을 논의한다.
- CNN-ILF와 CNN-BARC의 두 CNN 기반 도구를 갖춘 DLVC 사례 연구를 제시하고 압축 이득을 보고한다.
- 구성에 따라 HEVC와의 실험적 비교를 제공하여 잠재적 이득과 트레이드오프를 보여준다.
실험 결과
연구 질문
- RQ1이미지/비디오 코딩에 적용된 주요 딥러닝 패러다임은 무엇이며, 접근 방식과 목표에서 어떻게 다른가?
- RQ2딥 스킴과 딥 툴의 압축 이득 달성에서의 성능과 실용성은 어떻게 비교되는가?
- RQ3전통 코덱 내의 CNN 기반 도구가 DLVC의 속도-비트레이트 성능과 복잡도에 어떤 영향을 미치는가?
주요 결과
- 딥 스킴은 특정 경우 전통적인 이미지 코딩 효율성과 같거나 이를 초과할 수 있지만, 비디오 코딩 이득은 지금까지 HEVC에 비해 상대적으로 보수적이다.
- 전통 코덱에 내장된 딥 도구는 HEVC를 넘어서는 일부 측면에서 유망한 압축 성능을 보인다.
- CNN-ILF와 CNN-BARC를 포함한 DLVC 프로토타입은 HEVC 대비 평균 비트레이트 절감이 무작위 접근에서 39.6%, 저지연에서 33.0%에 달한다.
- 논문은 압축 효율성, 인코딩/디코딩 복잡도, 지각 품질, 보편적 적용성, 도구 통합 간의 트레이드오프를 강조한다.
- 향후 연구를 위한 DLVC 코드가 공개되었다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.