Skip to main content
QUICK REVIEW

[Paper Review] Opal: Multimodal Image Generation for News Illustration

Vivian Liu, Han Qiao|arXiv (Cornell University)|Apr 19, 2022
Multimodal Machine Learning Applications4 citations
TL;DR

Opal is a multimodal text-to-image generation system designed for news illustration that guides users through structured prompt engineering using article tone, keywords, and artistic styles. Leveraging GPT-3 for semantic suggestion and a three-stage pipeline, it enables users to generate two times more usable illustrations than without the system, significantly improving efficiency and creative output in co-creative workflows.

ABSTRACT

Advances in multimodal AI have presented people with powerful ways to create images from text. Recent work has shown that text-to-image generations are able to represent a broad range of subjects and artistic styles. However, finding the right visual language for text prompts is difficult. In this paper, we address this challenge with Opal, a system that produces text-to-image generations for news illustration. Given an article, Opal guides users through a structured search for visual concepts and provides a pipeline allowing users to generate illustrations based on an article's tone, keywords, and related artistic styles. Our evaluation shows that Opal efficiently generates diverse sets of news illustrations, visual assets, and concept ideas. Users with Opal generated two times more usable results than users without. We discuss how structured exploration can help users better understand the capabilities of human AI co-creative systems.

Motivation & Objective

  • To address the challenge of inconsistent and unpredictable text-to-image generation in news illustration by introducing a structured, guided workflow.
  • To reduce the trial-and-error burden in prompt engineering by leveraging large language models to suggest relevant keywords, tones, and artistic styles.
  • To improve the efficiency and quality of editorial image generation by integrating multimodal AI with human editorial judgment.
  • To evaluate whether structured exploration with LLM-powered suggestions enhances user performance in generating usable illustrations.
  • To explore how generative AI can augment—not replace—human illustrators in co-creative news design processes.

Proposed method

  • Opal employs a three-stage pipeline: (1) article input and keyword extraction using GPT-3, (2) tone and emotional characterization via NLP, and (3) artistic style suggestion using semantic search and LLM associations.
  • It uses GPT-3 as a knowledge base to generate keyword, tone, and style suggestions based on article content, enabling systematic prompt construction.
  • The system applies semantic search to map article concepts to related visual concepts and artistic styles, improving prompt relevance.
  • It supports users in generating image galleries through a text-based interface that structures exploration of subject, tone, and style.
  • The pipeline is designed to reduce stochasticity and improve consistency by guiding users toward high-quality, semantically aligned prompts.
  • User studies compare performance with and without Opal, measuring generation efficiency and usability of outputs.
Figure 1 . A screenshot of the Opal system, which helps users create news illustrations using a text-to-image generative AI model. The system here has generated a gallery of images for an article on ”climate change”. The participant is guided through the generation process with a structured pipeline
Figure 1 . A screenshot of the Opal system, which helps users create news illustrations using a text-to-image generative AI model. The system here has generated a gallery of images for an article on ”climate change”. The participant is guided through the generation process with a structured pipeline

Experimental results

Research questions

  • RQ1Can structured, LLM-guided prompt engineering improve the efficiency and quality of text-to-image generation for news illustration?
  • RQ2How do users perform in generating usable illustrations when supported by a system that suggests keywords, tones, and artistic styles?
  • RQ3To what extent can large language models like GPT-3 provide human-benchmark-quality suggestions for visual concepts with minimal user effort?
  • RQ4How does Opal support the co-creative process between human illustrators and generative AI in a real-world editorial context?
  • RQ5What are the limitations of text-based AI prompting in image generation, particularly when compared to traditional image-based or direct manipulation workflows?

Key findings

  • Users with Opal generated two times more usable illustrations than users without the system, demonstrating a significant improvement in output quality and efficiency.
  • The structured pipeline reduced the time and cognitive load associated with prompt engineering, enabling faster iteration and idea generation.
  • LLM-generated suggestions for keywords, tones, and styles achieved performance close to human benchmarks, requiring significantly less effort from users.
  • Participants reported that Opal’s AI-assisted suggestions enhanced their creative process, providing useful references, inspiration, and design material.
  • Despite the benefits, users expressed a preference for image-based prompting and direct manipulation, indicating a gap in current text-only interface design.
  • The study confirmed that generative AI should augment, not replace, human illustrators, as artistic judgment and conceptual understanding remain essential throughout the process.
Figure 2 . Text-to-image generations that were successful with news illustrators during the co-design process. These generations captured design patterns discussed in the formative study, where subjects and styles were suggested based on keywords and tones. For example, in the top left, ”glitch art”
Figure 2 . Text-to-image generations that were successful with news illustrators during the co-design process. These generations captured design patterns discussed in the formative study, where subjects and styles were suggested based on keywords and tones. For example, in the top left, ”glitch art”

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.