Skip to main content
QUICK REVIEW

[Paper Review] Visual Prompting in Multimodal Large Language Models: A Survey

Junda Wu, Zhehao Zhang|arXiv (Cornell University)|Sep 5, 2024
Multimodal Machine Learning Applications4 citations
TL;DR

This survey presents the first comprehensive analysis of visual prompting in multimodal large language models (MLLMs), categorizing visual prompt types, generative methods, and alignment techniques to enhance fine-grained visual grounding, compositional reasoning, and hallucination mitigation. It identifies visual prompting as a critical paradigm for improving MLLM perception and controllability through pixel-level, heterogeneous prompts and systematic training strategies.

ABSTRACT

Multimodal large language models (MLLMs) equip pre-trained large-language models (LLMs) with visual capabilities. While textual prompting in LLMs has been widely studied, visual prompting has emerged for more fine-grained and free-form visual instructions. This paper presents the first comprehensive survey on visual prompting methods in MLLMs, focusing on visual prompting, prompt generation, compositional reasoning, and prompt learning. We categorize existing visual prompts and discuss generative methods for automatic prompt annotations on the images. We also examine visual prompting methods that enable better alignment between visual encoders and backbone LLMs, concerning MLLM's visual grounding, object referring, and compositional reasoning abilities. In addition, we provide a summary of model training and in-context learning methods to improve MLLM's perception and understanding of visual prompts. This paper examines visual prompting methods developed in MLLMs and provides a vision of the future of these methods.

Motivation & Objective

  • To address the limitations of textual prompting in MLLMs, which often lead to inaccurate visual grounding and hallucinations.
  • To provide a systematic categorization of visual prompting techniques, including bounding boxes, markers, pixel-level prompts, and soft prompts.
  • To examine methods for automatic visual prompt generation and their integration into MLLM inference for improved perception and reasoning.
  • To analyze alignment techniques—via pre-training, fine-tuning, and in-context learning—that reduce misinterpretation of visual prompts and enhance compositional reasoning.
  • To explore the role of visual prompting in mitigating bias and enabling controllable visual generation in diffusion and restoration models.

Proposed method

  • Categorizes visual prompting into four main types: bounding boxes, markers, pixel-level prompts, and soft prompts, based on spatial granularity and modality.
  • Reviews generative methods for automatic visual prompt annotation, including detection-based, segmentation-based, and diffusion-based prompt generation.
  • Analyzes how visual prompts are integrated into MLLMs to improve visual grounding, object referring, and compositional reasoning through cross-modal attention mechanisms.
  • Examines model training strategies such as instruction tuning, contrastive learning, and in-context learning to align visual encoders with LLM backbones and reduce prompt misinterpretation.
  • Evaluates the impact of visual prompting on reducing hallucinations via mutual information decoding and joint visual-textual prompting.
  • Investigates the use of visual prompting in visual generation tasks, including text-to-image diffusion models, image editing, and 3D generation, using control mechanisms like ControlNet and T2I-Adapter.

Experimental results

Research questions

  • RQ1How can visual prompting improve the accuracy of visual grounding and object referring in MLLMs compared to textual prompting alone?
  • RQ2What are the key categories and forms of visual prompts, and how do they differ in spatial and functional granularity?
  • RQ3How can visual prompts be automatically generated for diverse images using detection, segmentation, or diffusion models?
  • RQ4What training and in-context learning strategies effectively align MLLMs with visual prompts to reduce hallucinations and language bias?
  • RQ5In what ways does visual prompting enhance compositional reasoning and controllability in multimodal generation tasks?

Key findings

  • Visual prompting significantly improves visual grounding and reduces hallucinations by enabling pixel-level, fine-grained instructions that are more aligned with visual input.
  • Joint use of visual and textual prompts reduces object hallucination in MLLMs, particularly in multi-object perception tasks.
  • Mutual information decoding strategies amplify the influence of visual prompts on generation, improving grounding and reducing hallucination.
  • Prompt generation via detection models enhances relevance and precision in visual prompting, especially when key concepts are extracted from textual prompts.
  • Visual prompting mitigates language bias by grounding model outputs in visual features, reducing reliance on spurious correlations.
  • In visual generation, visual prompting via ControlNet and T2I-Adapter enables spatial control and improves performance across six visual tasks after fine-tuning.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.