Skip to main content
QUICK REVIEW

[Paper Review] From Sora What We Can See: A Survey of Text-to-Video Generation

Rui Sun, Yumin Zhang|arXiv (Cornell University)|May 17, 2024
Video Analysis and Summarization4 citations
TL;DR

This survey provides a comprehensive analysis of text-to-video (T2V) generation by deconstructing OpenAI's Sora, categorizing literature across three dimensions—evolutionary generators, excellence pursuit (duration, resolution, quality), and realistic panoply (motion, complexity, layout). It reviews key models, datasets, metrics, and identifies critical challenges and future research directions in T2V generation.

ABSTRACT

With impressive achievements made, artificial intelligence is on the path forward to artificial general intelligence. Sora, developed by OpenAI, which is capable of minute-level world-simulative abilities can be considered as a milestone on this developmental path. However, despite its notable successes, Sora still encounters various obstacles that need to be resolved. In this survey, we embark from the perspective of disassembling Sora in text-to-video generation, and conducting a comprehensive review of literature, trying to answer the question, extit{From Sora What We Can See}. Specifically, after basic preliminaries regarding the general algorithms are introduced, the literature is categorized from three mutually perpendicular dimensions: evolutionary generators, excellent pursuit, and realistic panorama. Subsequently, the widely used datasets and metrics are organized in detail. Last but more importantly, we identify several challenges and open problems in this domain and propose potential future directions for research and development.

Motivation & Objective

  • To provide a systematic review of text-to-video (T2V) generation literature by analyzing OpenAI’s Sora as a benchmark for advanced video generation.
  • To categorize existing T2V methods along three orthogonal dimensions: evolutionary models, pursuit of video excellence (duration, resolution, quality), and realistic video generation (motion, complexity, layout).
  • To organize and evaluate widely used datasets and benchmark metrics in T2V research.
  • To identify persistent challenges and open problems in T2V generation, especially those highlighted by Sora’s limitations.
  • To propose future research directions grounded in Sora’s capabilities and shortcomings, supporting the advancement of artificial general intelligence through video generation.

Proposed method

  • The survey conducts a multi-dimensional literature review, classifying T2V methods based on generative model evolution (GAN/VAE, autoregressive, diffusion-based), excellence in video attributes (duration, resolution, quality), and realism (motion coherence, complex scenes, layout rationality).
  • It analyzes Sora’s architecture, particularly its use of the DiT (Diffusion Transformer) model, which replaces traditional U-Net and enables long-context, high-fidelity video generation.
  • The paper evaluates T2V datasets and metrics by source and domain, including benchmarks for quality, diversity, and temporal consistency.
  • It identifies key challenges such as motion inconsistency among multiple objects, lack of physical realism, and data sparsity in long video generation.
  • The survey draws insights from Sora’s world-simulative capabilities to propose future integration with digital twins and normative AI frameworks.
  • It proposes future research directions, including explainable AI, privacy-preserving generation, fairness, and robustness in generative video systems.

Experimental results

Research questions

  • RQ1How does Sora’s DiT-based architecture enable minute-long, high-resolution video generation, and what are its technical advantages over prior models?
  • RQ2What are the key dimensions that define excellence in text-to-video generation, and how do current models perform across duration, resolution, and visual quality?
  • RQ3What challenges remain in generating physically coherent and temporally consistent motion in complex, multi-object scenes?
  • RQ4How can text-to-video models be integrated into digital twin systems to improve simulation fidelity and real-time responsiveness?
  • RQ5What normative and ethical frameworks are necessary to ensure responsible development and deployment of advanced T2V models like Sora?

Key findings

  • Sora achieves minute-long video generation with high resolution and seamless quality, surpassing prior T2V models that were limited to short, low-resolution outputs.
  • The DiT (Diffusion Transformer) architecture enables Sora to process long text prompts and generate coherent, high-fidelity videos, outperforming traditional U-Net-based models.
  • Despite its strengths, Sora struggles with motion coherence among multiple objects, indicating a key open challenge in multi-object video generation.
  • The survey identifies a lack of standardized evaluation metrics for long-form video generation, particularly for temporal consistency and physical plausibility.
  • Sora’s capabilities suggest strong potential for integration with digital twin systems, especially in enhancing data completion and real-time simulation using physical priors.
  • The authors emphasize the urgent need for normative AI frameworks addressing explainability, privacy, fairness, and ethical use in generative video systems.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.