Skip to main content
QUICK REVIEW

[Paper Review] Multimodal Analysis Of Google Bard And GPT-Vision: Experiments In Visual Reasoning

David Noever, Samantha Elizabeth Miller Noever|arXiv (Cornell University)|Aug 17, 2023
Multimodal Machine Learning ApplicationsComputer Science3 citations
TL;DR

This study evaluates Google Bard and GPT-Vision through 64 visual reasoning tasks, revealing their strengths in visual CAPTCHA solving and visual text recognition, yet significant limitations in reconstructing ASCII art, analyzing structured grids like Tic-Tac-Toe, and forecasting next scenes—indicating over-reliance on visual guesswork rather than robust visual understanding.

ABSTRACT

Addressing the gap in understanding visual comprehension in Large Language Models (LLMs), we designed a challenge-response study, subjecting Google Bard and GPT-Vision to 64 visual tasks, spanning categories like "Visual Situational Reasoning" and "Next Scene Prediction." Previous models, such as GPT4, leaned heavily on optical character recognition tools like Tesseract, whereas Bard and GPT-Vision, akin to Google Lens and Visual API, employ deep learning techniques for visual text recognition. However, our findings spotlight both vision-language model's limitations: while proficient in solving visual CAPTCHAs that stump ChatGPT alone, it falters in recreating visual elements like ASCII art or analyzing Tic Tac Toe grids, suggesting an over-reliance on educated visual guesses. The prediction problem based on visual inputs appears particularly challenging with no common-sense guesses for next-scene forecasting based on current "next-token" multimodal models. This study provides experimental insights into the current capacities and areas for improvement in multimodal LLMs.

Motivation & Objective

  • To investigate the visual reasoning capabilities of recent multimodal LLMs, specifically Google Bard and GPT-Vision, in comparison to earlier models like GPT-4.
  • To identify the limitations of vision-language models in tasks requiring precise visual comprehension beyond optical character recognition.
  • To evaluate how well these models perform on tasks involving visual situational reasoning, next-scene prediction, and reconstruction of visual patterns such as ASCII art and game boards.
  • To assess whether the integration of deep learning-based visual understanding in models like Bard and GPT-Vision leads to improved reasoning over purely OCR-dependent models.

Proposed method

  • A challenge-response study was conducted, presenting 64 diverse visual reasoning tasks to Google Bard and GPT-Vision.
  • Tasks were grouped into categories including 'Visual Situational Reasoning' and 'Next Scene Prediction' to assess contextual and sequential understanding.
  • Visual inputs included CAPTCHAs, ASCII art, Tic-Tac-Toe grids, and scene transitions to test pattern recognition and structural reasoning.
  • The models were evaluated based on their ability to accurately interpret, describe, and reason about visual inputs without relying on external tools.
  • Performance was measured by comparing model outputs against ground-truth answers, focusing on correctness and coherence.
  • The study contrasted results with prior models like GPT-4, which relied on external OCR tools such as Tesseract, to highlight shifts in architectural approach.

Experimental results

Research questions

  • RQ1How do Google Bard and GPT-Vision perform on visual reasoning tasks that require understanding of visual structure and context?
  • RQ2To what extent do these models rely on visual guesswork rather than accurate visual comprehension in tasks like ASCII art reconstruction or grid analysis?
  • RQ3Can these models effectively predict the next scene in a visual sequence based on current multimodal inputs?
  • RQ4How do their visual reasoning capabilities compare to earlier models that used external OCR tools like Tesseract?
  • RQ5What are the specific failure modes in visual reasoning when models are confronted with tasks requiring precise spatial or structural reasoning?

Key findings

  • Google Bard and GPT-Vision successfully solved visual CAPTCHAs that were unsolvable by ChatGPT alone, demonstrating improved visual perception over earlier models.
  • The models struggled significantly with reconstructing ASCII art, indicating a lack of precise understanding of visual patterns and spatial arrangements.
  • Performance on Tic-Tac-Toe grid analysis was poor, revealing limitations in recognizing and reasoning about structured, symbolic visual layouts.
  • Next-scene prediction tasks showed no reliable common-sense reasoning, with models frequently generating incorrect or incoherent continuations.
  • Despite using deep learning for visual text recognition, the models exhibited an over-reliance on educated visual guesses rather than accurate visual comprehension.
  • The study confirms that current multimodal LLMs still lack robust visual reasoning capabilities, especially in tasks requiring structural or contextual inference.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.