Skip to main content
QUICK REVIEW

[논문 리뷰] Measuring the Success of Diffusion Models at Imitating Human Artists

Stephen Casper, Zifan Guo|arXiv (Cornell University)|2023. 07. 08.
Generative Adversarial Networks and Image SynthesisComputer Science인용 수 3
한 줄 요약

이 논문은 스타일이 특정 인간 예술가를 모방하는 데 얼마나 잘 맞는지 정량적으로 측정할 수 있는 제로샷, CLIP 기반 방법을 제안한다. 스타일이 특정 예술가의 것임을 나타내는 프롬프트를 사용해 이미지를 생성하고, 그 결과를 CLIP로 분류함으로써, 저자들은 스타일러니케이션 모델이 81.0%의 경우에서 예술가를 정확히 식별함을 보여주었다. 이는 개별 예술가의 스타일을 넓게 모방할 수 있음을 시사한다.

ABSTRACT

Modern diffusion models have set the state-of-the-art in AI image generation. Their success is due, in part, to training on Internet-scale data which often includes copyrighted work. This prompts questions about the extent to which these models learn from, imitate, or copy the work of human artists. This work suggests that tying copyright liability to the capabilities of the model may be useful given the evolving ecosystem of generative models. Specifically, much of the legal analysis of copyright and generative systems focuses on the use of protected data for training. As a result, the connections between data, training, and the system are often obscured. In our approach, we consider simple image classification techniques to measure a model's ability to imitate specific artists. Specifically, we use Contrastive Language-Image Pretrained (CLIP) encoders to classify images in a zero-shot fashion. Our process first prompts a model to imitate a specific artist. Then, we test whether CLIP can be used to reclassify the artist (or the artist's work) from the imitation. If these tests match the imitation back to the original artist, this suggests the model can imitate that artist's expression. Our approach is simple and quantitative. Furthermore, it uses standard techniques and does not require additional training. We demonstrate our approach with an audit of Stable Diffusion's capacity to imitate 70 professional digital artists with copyrighted work online. When Stable Diffusion is prompted to imitate an artist from this set, we find that the artist can be identified from the imitation with an average accuracy of 81.0%. Finally, we also show that a sample of the artist's work can be matched to these imitation images with a high degree of statistical reliability. Overall, these results suggest that Stable Diffusion is broadly successful at imitating individual human artists.

연구 동기 및 목표

  • 확산 모델이 특정 인간 예술가를 얼마나 잘 모방하는지 평가하기 위한 실용적이고 정량적인 방법을 개발하기 위해.
  • AI 생성 예술의 저작권 책임을 평가할 때 학습 데이터가 아닌 모델의 능력에 초점을 맞추기 위해.
  • 기본적인 기계학습 기법이 AI 모방에 관한 법적 주장을 분석하는 데 사용될 수 있음을 보여주기 위해.
  • 스테이블 디퓨전이 개별 디지털 예술가의 시각적 스타일을 어느 정도 재현할 수 있는지 평가하기 위해.
  • 사용 가능한 모델과 표준 기법을 활용해 AI 모방 성공 여부를 재현 가능한 벤치마크로 제공하기 위해.

제안 방법

  • 스테이블 디퓨전 v1.5를 사용해 각 예술가당 10장의 이미지를 프롬프트 'Artwork from <artist name>' 형식으로 생성
  • 생성된 이미지와 텍스트 레이블(예상 예술가 이름 및 기본 레이블 포함)을 CLIP의 이미지 및 텍스트 인코더를 사용해 인코딩
  • 이미지 및 텍스트 임베딩 간 코사인 유사도를 계산하여 각 생성된 이미지를 분류하고, 유사도가 가장 높은 레이블을 선택
  • 각 예술가당 분류 실험을 10회 반복하여 분산을 줄이고 평균 정확도를 계산
  • 두 가지 베이스라인과 결과를 비교: 무작위 예술가 이름(8.6% 정확도)과 무작위 추측(1.4% 정확도)
  • 두 번째 실험으로 실제 예술가의 작품과 생성된 모의물 간의 일치를 CLIP 임베딩과 통계적 유의성 검정(보프에르니 보정을 적용한 순위합 검정)을 사용해 수행
Figure 1: Identifying human artists from Stable Diffusion Imitations. For each artist, we generate an imitation image from Stable Diffusion with the prompt “Artwork from $<$ artist name $>$ .” Next, we encode the image with a CLIP image encoder (Radford et al., 2021 ) . We also encode labels corresp
Figure 1: Identifying human artists from Stable Diffusion Imitations. For each artist, we generate an imitation image from Stable Diffusion with the prompt “Artwork from $<$ artist name $>$ .” Next, we encode the image with a CLIP image encoder (Radford et al., 2021 ) . We also encode labels corresp

실험 결과

연구 질문

  • RQ1확산 모델, 예를 들어 스타일러니케이션 모델이 인간 예술가의 시각적 스타일을 어느 정도 모방할 수 있는가?
  • RQ2표준 제로샷 이미지 분류 기법, 예를 들어 CLIP은 AI로 생성된 이미지를 특정 예술가로 식별하고 소속시킬 수 있는가?
  • RQ3모델의 모방 능력은 무작위 확률이나 기준 모델에 비해 어떻게 비교되는가?
  • RQ4AI로 생성된 모의물과 실제 예술가의 작품 간 유사성은 통계적으로 유의미한가?
  • RQ5학습 데이터에 의존하지 않고 모델의 출력 행동에 초점을 맞춰 모방 성공 여부를 독립적으로 측정할 수 있는가?

주요 결과

  • 스테이블 디퓨전은 70명의 전문 디지털 예술가에 대해 생성된 이미지의 예술가를 평균 81.0%의 정확도로 올바르게 분류한다.
  • 70명의 예술가 중 69명이 대부분의 시도에서 올바른 예술가를 식별함으로써, 넓고 일관된 모방 능력을 보여준다.
  • 무작위 이름 기반 베이스라인은 오직 8.6%의 정확도를 기록했고, 무작위 추측은 1.4%에 그쳐 결과가 우연이 아님을 확인한다.
  • 실제 작품과 생성된 모의물 간의 매칭을 수행한 결과, 70명 중 63명(90%)이 통계적으로 유의미한 유사성(p < 0.05, 보프에르니 보정 후)을 보였다.
  • 다른 예술가 집합에 대해서도 결과가 강인함을 확인했으며, LAION 데이터셋에서 가장 많이 등장하는 250명의 예술가에 대해 유사한 81.2%의 정확도를 관찰했다.
  • 이 방법은 모의물이 목표 예술가의 실제 작품보다 다른 예술가의 작품보다 더 유사하다는 것을 성공적으로 식별함으로써, 모방의 정밀도를 확인한다.
Figure 2: Example images generated by Stable Diffusion from prompts of the form “Artwork from $<$ artist’s name $>$ ”. Using the method depicted in Figure 1 , we show that the artists used in the prompts can often be classified from these imitations of their work.
Figure 2: Example images generated by Stable Diffusion from prompts of the form “Artwork from $<$ artist’s name $>$ ”. Using the method depicted in Figure 1 , we show that the artists used in the prompts can often be classified from these imitations of their work.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.