Skip to main content
QUICK REVIEW

[Paper Review] Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Robotics Team, Petko Georgiev|arXiv (Cornell University)|Mar 8, 2024
Semantic Web and OntologiesComputer Science276 citations
TL;DR

Gemini 1.5 introduces two long-context multimodal models (Gemini 1.5 Pro and Gemini 1.5 Flash) that recall and reason over millions of tokens, achieving near-perfect long-context retrieval and state-of-the-art performance on long-document QA, long-video QA, and long-context ASR.

ABSTRACT

In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio. The family includes two new models: (1) an updated Gemini 1.5 Pro, which exceeds the February version on the great majority of capabilities and benchmarks; (2) Gemini 1.5 Flash, a more lightweight variant designed for efficiency with minimal regression in quality. Gemini 1.5 models achieve near-perfect recall on long-context retrieval tasks across modalities, improve the state-of-the-art in long-document QA, long-video QA and long-context ASR, and match or surpass Gemini 1.0 Ultra's state-of-the-art performance across a broad set of benchmarks. Studying the limits of Gemini 1.5's long-context ability, we find continued improvement in next-token prediction and near-perfect retrieval (>99%) up to at least 10M tokens, a generational leap over existing models such as Claude 3.0 (200k) and GPT-4 Turbo (128k). Finally, we highlight real-world use cases, such as Gemini 1.5 collaborating with professionals on completing their tasks achieving 26 to 75% time savings across 10 different job categories, as well as surprising new capabilities of large language models at the frontier; when given a grammar manual for Kalamang, a language with fewer than 200 speakers worldwide, the model learns to translate English to Kalamang at a similar level to a person who learned from the same content.

Motivation & Objective

  • Advance multimodal understanding with extremely long context windows (millions of tokens).
  • Provide compute-efficient variants that preserve quality (Gemini 1.5 Pro and Gemini 1.5 Flash).
  • Demonstrate improvements in long-context retrieval, long-document QA, long-video QA, and long-context ASR.
  • Show practical real-world impact and surprising capabilities in low-resource language tasks.

Proposed method

  • Develop two models: Gemini 1.5 Pro (improved over February version across benchmarks) and Gemini 1.5 Flash (more efficient with minimal quality loss).
  • Demonstrate near-perfect retrieval (>99%) up to 10 million tokens across modalities.
  • Evaluate on long-document QA, long-video QA, and long-context ASR benchmarks against prior models including Gemini 1.0 Ultra.
  • Analyze next-token prediction performance as context length scales to assess long-context limits.
  • Present real-world use cases illustrating time savings and cross-domain capabilities.

Experimental results

Research questions

  • RQ1How well can Gemini 1.5 recall and reason over millions of tokens across text, video, and audio?
  • RQ2What are the trade-offs between accuracy and efficiency in Gemini 1.5 Pro versus Gemini 1.5 Flash?
  • RQ3Do long-context models achieve state-of-the-art performance on long-document QA, long-video QA, and long-context ASR?
  • RQ4What are the practical real-world impacts and limitations of deploying Gemini 1.5 in diverse tasks (including low-resource languages)?

Key findings

  • Gemini 1.5 achieves near-perfect retrieval (>99%) for up to 10M tokens.
  • Gemini 1.5 Pro outperforms the February version on most capabilities and benchmarks.
  • Gemini 1.5 Flash offers efficiency with minimal regression in quality compared to Pro.
  • The models set new state-of-the-art results for long-document QA, long-video QA, and long-context ASR.
  • In real-world scenarios, Gemini 1.5 enables 26–75% time savings across 10 job categories.
  • The models demonstrate surprising capabilities, such as learning to translate Kalamang from grammar material at a level comparable to a learner with the same content.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.