Skip to main content
QUICK REVIEW

[논문 리뷰] Quality-aware Pre-trained Models for Blind Image Quality Assessment

Kai Zhao, Kun Yuan|arXiv (Cornell University)|2023. 03. 01.
Image and Video Quality Assessment인용 수 6
한 줄 요약

이 논문은 레이블이 없는 품질 데이터 없이 ImageNet에서 품질 민감한 표현을 학습하는 자기지도 대비 학습 프레임워크를 사용하여 빌드된 품질 인식 사전 훈련(QPT) 모델을 제안한다. 2×10⁷가 넘는 가능한 왜곡을 포함하는 복잡한 왜곡 공간과 품질 인식 대비 손실을 설계함으로써, QPT는 다섯 가지 벤치마크에서 후행 BIQA 성능을 크게 향상시켜 SRCC 기준 최대 8.75% 향상된다.

ABSTRACT

Blind image quality assessment (BIQA) aims to automatically evaluate the perceived quality of a single image, whose performance has been improved by deep learning-based methods in recent years. However, the paucity of labeled data somewhat restrains deep learning-based BIQA methods from unleashing their full potential. In this paper, we propose to solve the problem by a pretext task customized for BIQA in a self-supervised learning manner, which enables learning representations from orders of magnitude more data. To constrain the learning process, we propose a quality-aware contrastive loss based on a simple assumption: the quality of patches from a distorted image should be similar, but vary from patches from the same image with different degradations and patches from different images. Further, we improve the existing degradation process and form a degradation space with the size of roughly $2 imes10^7$. After pre-trained on ImageNet using our method, models are more sensitive to image quality and perform significantly better on downstream BIQA tasks. Experimental results show that our method obtains remarkable improvements on popular BIQA datasets.

연구 동기 및 목표

  • 딥 러닝 성능을 저해하는 빌드 인 품질 평가(BIQA)에서 레이블이 부족한 데이터의 한계를 해결한다.
  • 의미보다는 이미지 품질에 중점을 둔 기존 사전 훈련 모델(예: ImageNet)의 열악한 전이 가능성 문제를 해결한다.
  • 저수준의 왜곡과 콘텐츠-품질 상호작용에 민감한 BIQA에 적합한 자기지도 사전 과제를 개발한다.
  • 실제 이미지 왜곡을 시뮬레이션하고 표현 학습을 향상시키기 위해 대규모이고 현실적인 왜곡 공간을 구축한다.

제안 방법

  • 실제 왜곡을 시뮬레이션하기 위해 순서 뒤섞기, 고차원 연산, 스킵 연결을 통합한 새로운 왜곡 프로세스를 설계한다.
  • 약 2×10⁷개의 고유한 왜곡 유형을 포함하는 왜곡 공간을 구축하여 대비 학습을 위한 다양한 데이터 증강을 가능하게 한다.
  • 동일한 왜곡 이미지에서 추출한 패치를 양성 쌍으로 간주하고, 서로 다른 이미지 간의 패치(콘텐츠 기반 음성)와 서로 다른 왜곡 간의 패치(왜곡 기반 음성)를 구별하는 품질 인식 대비 손실(QC-Loss)을 제안한다.
  • MoCoV2 프레임워크를 변형하여 자기지도 사전 훈련을 수행하고, QC-Loss를 사용해 모델이 의미적 콘텐츠보다 이미지 품질에 민감하도록 훈련한다.
  • 제안된 QC-Loss와 왜곡 증강을 사용해 ImageNet에서 사전 훈련을 수행한 후, 후행 BIQA 데이터셋에서 미세조정 또는 선형 프로빙을 수행한다.
  • 동일한 사전 훈련 가중치를 여러 BIQA 벤치마크에 그대로 재사용함으로써 일반화 능력을 확보한다.
Figure 1 : The two images in the first row are sampled from BIQA dataset CLIVE [ 20 ] . Although they have the same semantic meaning, their perceptual qualities are quite different: their mean opinion scores (MOS) are 31.83 and 86.35. The second row shows modified versions of the first two images an
Figure 1 : The two images in the first row are sampled from BIQA dataset CLIVE [ 20 ] . Although they have the same semantic meaning, their perceptual qualities are quite different: their mean opinion scores (MOS) are 31.83 and 86.35. The second row shows modified versions of the first two images an

실험 결과

연구 질문

  • RQ1이미지 품질 차이에 중점을 둔 자기지도 사전 과제가 빌드 인 품질 평가를 위한 표현 학습을 향상시킬 수 있는가?
  • RQ2크고 다양한 왜곡 공간은 BIQA에서 사전 훈련된 특징의 품질 인식 능력에 어떤 영향을 미치는가?
  • RQ3대비 학습 목표에서 콘텐츠 기반 음성 샘플과 왜곡 기반 음성 샘플의 상대 기여도는 어떠한가?
  • RQ4기본 대비 학습 또는 감독 사전 훈련과 비교해 제안된 QC-Loss는 후행 BIQA 성능을 얼마나 향상시키는가?

주요 결과

  • 제안된 QPT 방법은 다섯 가지 표준 BIQA 벤치마크에서 최신 기술 성능을 달성했으며, BID 데이터셋에서 SRCC 0.8875, PLCC 0.9109를 기록했다.
  • CLIVE 데이터셋에서 QPT는 SRCC 0.8947, PLCC 0.9141을 달성했으며, 선형 프로빙 평가에서 MoCo보다 SRCC 기준 8.75%, PLCC 기준 8.17% 향상되었다.
  • 절단 실험 결과, 상위 샘플(콘텐츠 기반)과 내부 샘플(왜곡 기반) 음성 쌍을 모두 포함할 경우 성능이 가장 우수하며, 상위 샘플 음성 쌍을 제거할 경우 성능이 크게 떨어지는 것으로 확인되었다.
  • 왜곡 프로세스에서 순서 뒤섞기, 고차원 연산, 스킵 연결을 모두 조합할 경우 성능이 가장 높았으며, BID에서 SRCC 0.8875, PLCC 0.9109를 기록했다.
  • QPT 가중치를 사용한 엔드 투 엔드 미세조정이 가장 높은 성능을 기록했으며(BID에서 SRCC: 0.8875, PLCC: 0.9109) 이는 선형 프로빙을 넘어서 학습된 특징의 효과성을 확인한다.
  • QPT 사전 훈련 모델은 다양한 데이터셋 간에 잘 일반화되며, 재훈련이나 적응 없이도 강력한 전이 능력을 보였다.
Figure 2 : Illustration of generating distorted images using different compositions of degradation. Compared with the process of fixed sequence, the introduced skip, shuffle and high-order largely increase the degradation space, covering diverse and realistic distortions.
Figure 2 : Illustration of generating distorted images using different compositions of degradation. Compared with the process of fixed sequence, the introduced skip, shuffle and high-order largely increase the degradation space, covering diverse and realistic distortions.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.