[논문 리뷰] User-Guided Personalized Image Aesthetic Assessment based on Deep Reinforcement Learning
이 논문은 사용자 가이드된 인터랙티브 이미지 토닝 및 랭킹을 통해 개인의 선호도를 모델링하기 위해 딥 강화학습(DRL)을 사용하는 사용자 지향 개인화된 이미지 미적 평가 프레임워크를 제안한다. 사용자 피드백에 기반해 반복적으로 향상 및 랭킹 정책 네트워크를 훈련시킴으로써, 정확한 개인화된 미적 분포를 생성하며, AVA 및 FLICKR-AES 데이터셋에서 최신 기술 수준의 성능을 달성한다.
Personalized image aesthetic assessment (PIAA) has recently become a hot topic due to its usefulness in a wide variety of applications such as photography, film and television, e-commerce, fashion design and so on. This task is more seriously affected by subjective factors and samples provided by users. In order to acquire precise personalized aesthetic distribution by small amount of samples, we propose a novel user-guided personalized image aesthetic assessment framework. This framework leverages user interactions to retouch and rank images for aesthetic assessment based on deep reinforcement learning (DRL), and generates personalized aesthetic distribution that is more in line with the aesthetic preferences of different users. It mainly consists of two stages. In the first stage, personalized aesthetic ranking is generated by interactive image enhancement and manual ranking, meanwhile two policy networks will be trained. The images will be pushed to the user for manual retouching and simultaneously to the enhancement policy network. The enhancement network utilizes the manual retouching results as the optimization goals of DRL. After that, the ranking process performs the similar operations like the retouching mentioned before. These two networks will be trained iteratively and alternatively to help to complete the final personalized aesthetic assessment automatically. In the second stage, these modified images are labeled with aesthetic attributes by one style-specific classifier, and then the personalized aesthetic distribution is generated based on the multiple aesthetic attributes of these images, which conforms to the aesthetic preference of users better.
연구 동기 및 목표
- 제한된 사용자 상호작용으로 인해 매우 주관적인 사용자 맞춤형 미적 선호도를 모델링하는 데 도전하는 것.
- 딥 강화학습 파이프라인에 인터랙티브 이미지 강화 및 랭킹을 통합하여 개인화된 이미지 미적 평가를 향상시키는 것.
- 최소한의 사용자 피드백을 사용하여 더 정확하고 사용자와 일치하는 미적 분포를 생성하는 것.
- 최적화된 상호작용 설계를 통해 사용자 피로도와 상호작용 시간을 줄이면서도 높은 평가 정확도를 유지하는 것.
제안 방법
- 프레임워크는 두 단계로 구성된다: 사용자 가이드된 이미지 미적 랭킹 및 개인화된 미적 분포 생성.
- 첫 번째 단계에서 사용자는 이미지를 인터랙티브하게 토닝하고 랭킹하며, 시스템은 이러한 행동을 바탕으로 딥 강화학습 기반의 강화 정책 네트워크를 훈련시킨다.
- 동일한 피드백을 사용하여 별도의 랭킹 정책 네트워크를 동시에 훈련시키며, 사용자 선호도에 기반한 이미지 순서 정렬을 최적화한다.
- 두 정책 네트워크는 반복적이고 번갈아가며 훈련되어 강화 및 랭킹 예측을 개선한다.
- 두 번째 단계에서 스타일별 분류기가 수정된 이미지에 대해 다중 미적 특성으로 레이블을 붙여 개인화된 미적 분포를 생성한다.
- 최종적으로, 새로운 이미지를 학습된 사용자 맞춤형 미적 분포와 비교하여 개인화된 평가를 유추한다.
실험 결과
연구 질문
- RQ1사용자 가이드된 이미지 강화 및 랭킹은 개인화된 미적 평가 정확도를 어떻게 향상시키는가?
- RQ2상호작용당 이미지 수와 총 상호작용 라운드 수 사이의 최적의 균형은 무엇인가?
- RQ3딥 강화학습은 최소한의 사용자 피드백을 사용하여 주관적인 사용자 선호도를 효과적으로 모델링할 수 있는가?
- RQ4인터랙티브 토닝의 통합은 예측된 사용자 미적 선호도와 실제 사용자 선호도 간의 일치도를 어떻게 향상시키는가?
주요 결과
- 제안된 방법은 개인화된 이미지 미적 평가에서 AVA 및 FLICKR-AES 데이터셋 모두에서 최신 기술 수준의 성능을 달성했다.
- 상호작용당 5장의 이미지와 총 5회의 상호작용 라운드를 사용할 경우, AVA 데이터셋에서 랭킹 상관관계가 0.6926에 도달했다.
- 상호작용당 평균 시간은 이미지 수에 따라 선형적으로 증가하여 15장일 경우 6.81분에 이를 것으로 나타나, 효율성과 인지 부하 사이의 상충 관계를 보였다.
- 연구 결과, 상호작용당 작은 수의 이미지(예: 5장)를 사용하고 여러 라운드를 거치는 것이 큰 배치에 비해 더 높은 정확도를 달성하고 사용자 피로도를 줄이는 데 유리하다는 것이 확인되었다.
- 프레임워크는 실제 사용자 선호도와 매우 유사한 개인화된 미적 분포를 성공적으로 생성하였으며, 사용자 랭킹 예측에서 높은 스피어만의 rho 값으로 검증되었다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.