Skip to main content
QUICK REVIEW

[Paper Review] The New Data and New Challenges in Multimedia Research.

Bart Thomée, David A. Shamma|arXiv (Cornell University)|Mar 5, 2015
Advanced Image and Video Retrieval Techniques18 references249 citations
TL;DR

The paper introduces the Yahoo Flickr Creative Commons 100 Million Dataset (YFCC100M), a publicly available multimedia collection of 100 million photos and videos with Creative Commons licensing, spanning from 2004 to 2014. It provides rich metadata and enables large-scale multimedia research, offering new challenges and opportunities in content understanding, representation, and sharing patterns.

ABSTRACT

We present the Yahoo Flickr Creative Commons 100 Million Dataset (YFCC100M), the largest public multimedia collection that has ever been released. The dataset contains a total of 100 million media objects, of which approximately 99.2 million are photos and 0.8 million are videos, all of which carry a Creative Commons license. Each media object in the dataset is represented by several pieces of metadata, e.g. Flickr identifier, owner name, camera, title, tags, geo, media source. The collection provides a comprehensive snapshot of how photos and videos were taken, described, and shared over the years, from the inception of Flickr in 2004 until early 2014. In this article we explain the rationale behind its creation, as well as the implications the dataset has for science, research, engineering, and development. We further present several new challenges in multimedia research that can now be expanded upon with our dataset.

Motivation & Objective

  • To create the largest publicly available multimedia dataset for research, enabling large-scale studies in multimedia understanding and content sharing.
  • To provide a comprehensive, long-term snapshot of user-generated photo and video content from Flickr's inception to 2014.
  • To support scientific and engineering advancements by offering a scalable, diverse, and well-annotated dataset with standardized metadata.
  • To identify and frame new challenges in multimedia research that emerge from analyzing such a large-scale, real-world dataset.

Proposed method

  • Collection of 100 million media objects from Flickr, including 99.2 million photos and 0.8 million videos, all under Creative Commons licensing.
  • Extraction and structuring of rich metadata per media object, including Flickr ID, owner name, camera, title, tags, geolocation, and media source.
  • Aggregation of data from Flickr's public API and database dumps, covering the period from 2004 to early 2014.
  • Design of a standardized data schema to ensure consistency and usability across diverse research applications.
  • Publication of the dataset as a public resource to support reproducible research and community-driven innovation.
  • Identification of emerging research challenges based on the dataset’s scale, diversity, and metadata richness.

Experimental results

Research questions

  • RQ1How can large-scale, real-world multimedia data be effectively collected and structured for broad research use?
  • RQ2What new challenges in multimedia understanding and content representation arise from analyzing 100 million user-generated media objects?
  • RQ3How do metadata such as tags, geolocation, and user-provided titles reflect human perception and content description patterns?
  • RQ4What insights can be gained about long-term trends in photo and video sharing behavior from 2004 to 2014?
  • RQ5How can public, licensed multimedia datasets enable scalable and reproducible research in computer vision and multimedia systems?

Key findings

  • The YFCC100M dataset comprises 100 million media objects, with 99.2 million photos and 0.8 million videos, all under Creative Commons licensing.
  • The dataset provides a comprehensive, long-term view of user-generated content from 2004 to early 2014, capturing evolving sharing and description behaviors.
  • Each media object is enriched with multiple metadata fields, including title, tags, geolocation, camera, and owner information, enabling rich analysis.
  • The dataset enables new research challenges in multimedia understanding, such as cross-modal retrieval, visual-semantic embedding, and content bias detection.
  • The availability of such a large-scale, public, and well-structured dataset opens new avenues for scalable and reproducible research in multimedia systems.
  • The dataset serves as a foundation for advancing research in computer vision, natural language processing, and social media analytics.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.