Skip to main content
QUICK REVIEW

[논문 리뷰] I Spy a Metaphor: Large Language Models and Diffusion Models Co-Create Visual Metaphors

Tuhin Chakrabarty, Arkadiy Saakyan|arXiv (Cornell University)|2023. 05. 24.
Multimodal Machine Learning Applications인용 수 4
한 줄 요약

이 논문은 체인 오브 쓰ought(Chain-of-Thought) 프롬프팅을 사용하여 언어적 은유의 구체적인 시각적 묘사를 생성하고, 이를 텍스트-이미지 디퓨전 모델의 입력으로 사용함으로써 대규모 언어 모델(Large Language Models, LLMs)과 디퓨전 모델이 협업하여 시각적 은유를 공동 생성하는 프레임워크를 제안한다. 주요 기여는 인간 검증이 완료된 고품질의 6,476개의 시각적 은유로 구성된 데이터셋(HAIVMet)으로, 이는 시각-언어 모델을 미세조정할 때 시각적 함의 작업에서 약 23점의 정확도 향상을 가능하게 한다.

ABSTRACT

Visual metaphors are powerful rhetorical devices used to persuade or communicate creative ideas through images. Similar to linguistic metaphors, they convey meaning implicitly through symbolism and juxtaposition of the symbols. We propose a new task of generating visual metaphors from linguistic metaphors. This is a challenging task for diffusion-based text-to-image models, such as DALL$\cdot$E 2, since it requires the ability to model implicit meaning and compositionality. We propose to solve the task through the collaboration between Large Language Models (LLMs) and Diffusion Models: Instruct GPT-3 (davinci-002) with Chain-of-Thought prompting generates text that represents a visual elaboration of the linguistic metaphor containing the implicit meaning and relevant objects, which is then used as input to the diffusion-based text-to-image models.Using a human-AI collaboration framework, where humans interact both with the LLM and the top-performing diffusion model, we create a high-quality dataset containing 6,476 visual metaphors for 1,540 linguistic metaphors and their associated visual elaborations. Evaluation by professional illustrators shows the promise of LLM-Diffusion Model collaboration for this task . To evaluate the utility of our Human-AI collaboration framework and the quality of our dataset, we perform both an intrinsic human-based evaluation and an extrinsic evaluation using visual entailment as a downstream task.

연구 동기 및 목표

  • 언어적 은유에서 시각적 은유를 생성하는 데 있어 암묵적인 의미와 조합적 구조를 모델링해야 하는 과제를 해결하기 위해.
  • 디퓨전 기반 텍스트-이미지 모델의 성능을 향상시켜, 단순한 해석을 초월해 비유적이고 상징적인 이미지를 생성할 수 있도록 하기 위해.
  • 미래의 시각적 은유 생성 및 시각-언어 이해 연구를 지원하기 위해 고품질의 인간 검증이 완료된 시각적 은유 데이터셋을 구축하기 위해.
  • LLM-디퓨전 협업 및 인간-AI 공동 창작의 효과성을 평가하기 위해, 의미 있고 조합적으로 정확한 시각적 은유를 생성하는 데에 초점을 맞추기 위해.

제안 방법

  • 언어적 은유의 암묵적 의미와 핵심 대상들을 포착하는 구체적인 시각적 묘사를 생성하기 위해 체인 오브 쓰ought 프롬프팅을 사용한 Instruct GPT-3(davinci-002)를 활용한다.
  • LLM가 생성한 시각적 묘사를 디퓨전 기반 텍스트-이미지 모델(DALL·E 2 또는 Stable Diffusion)의 프롬프트로 사용하여 최종 시각적 은유를 생성한다.
  • 인간 일러스트레이터가 LLM-디퓨전 파이프라인의 출력을 평가하고 보완함으로써 품질과 정확도를 확보하는 인간-AI 협업 프레임워크를 구현한다.
  • 1,540개의 언어적 은유에서 유래한 6,476개의 시각적 은유로 구성된 HAIVMet 데이터셋을 구축한다.
  • 전문 일러스트레이터의 인트라식 평가(내재적 평가)와 시각적 함의 최종 작업을 통한 외재적 평가를 통해 프레임워크를 평가한다.
  • HAIVMet 데이터셋에 기반해 시각-언어 모델(OFA-base)을 미세조정하고, SNLI-VE에만 기반해 미세조정한 결과와 성능을 비교한다.
Figure 1: Visual metaphors generated by DALL $\cdot$ E 2 for the linguistic metaphor “My bedroom is a pig sty".We can take the original verbal metaphor as the input (left) or use GPT-3 with Chain of Thought prompting (right).
Figure 1: Visual metaphors generated by DALL $\cdot$ E 2 for the linguistic metaphor “My bedroom is a pig sty".We can take the original verbal metaphor as the input (left) or use GPT-3 with Chain of Thought prompting (right).

실험 결과

연구 질문

  • RQ1LLM에서 체인 오브 쓰ought 프롬프팅이 디퓨전 모델이 생성하는 시각적 은유의 품질을 크게 향상시킬 수 있는가?
  • RQ2인간-AI 협업은 생성된 시각적 은유의 조합적 정확도와 비유적 충실도를 어떻게 향상시키는가?
  • RQ3HAIVMet 데이터셋에 기반해 미세조정할 경우, 기존 기준 베이스라인과 비교해 시각적 함의 작업 성능이 얼마나 향상되는가?
  • RQ4현재의 디퓨전 모델은 시각적 은유의 암묵적이고 비유적인 의미를 얼마나 잘 포착할 수 있는가?

주요 결과

  • LLM-디퓨전 협업 프레임워크는 시각적 은유 생성 품질을 크게 향상시키며, 전문 일러스트레이터가 LLM이 생성한 시각적 묘사를 기반으로 한 출력을 직접적인 언어적 은유 프롬프트보다 선호한다.
  • LLM가 생성한 묘사를 프롬프트로 사용했을 때 DALL·E 2는 Stable Diffusion v2.1보다 정확하고 개념적으로 일관된 시각적 은유를 더 잘 생성한다.
  • HAIVMet 데이터셋에 기반해 시각-언어 모델을 미세조정하면, SNLI-VE에만 기반해 미세조정한 경우에 비해 약 23점의 시각적 함의 정확도 향상이 이루어진다.
  • HAIVMet 데이터셋은 강력한 조합 일반화 능력을 보이며, '사랑은 욕망의 강 속에 있는 악어다'와 같은 복잡한 은유를 통합된 시각적 요소로 성공적으로 포착한다.
  • 인간-AI 협업은 텍스트-이미지 생성에서 부족한 명시성과 속성-객체 바인딩 문제를 해결하기 때문에 고품질 출력을 유지하는 데 필수적이다.
Figure 2: Human-AI collaboration framework (LLMs-Diffusion Model-Humans). Instruct GPT-3 with CoT prompting generates visual elaborations from linguistic metaphors, which are then validated and possibly edited by humans, if necessary. Visual elaborations are then used as input to DALL $\cdot$ E 2 to
Figure 2: Human-AI collaboration framework (LLMs-Diffusion Model-Humans). Instruct GPT-3 with CoT prompting generates visual elaborations from linguistic metaphors, which are then validated and possibly edited by humans, if necessary. Visual elaborations are then used as input to DALL $\cdot$ E 2 to

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.