[Paper Review] The YLI-MED Corpus: Characteristics, Procedures, and Plans
The YLI-MED corpus is a publicly available, human-annotated dataset of 50,700 user-generated videos from YFCC100M, with 2,000 videos labeled for 10 target multimedia events (e.g., birthday parties, weddings) and 48,700 non-event videos. It provides detailed annotations including event type, annotator agreement, confidence scores, and non-event attributes such as language and music, enabling training and evaluation of multimedia event detection systems with standardized training/test splits and feature bundles for audio, visual, and motion analysis.
The YLI Multimedia Event Detection corpus is a public-domain index of videos with annotations and computed features, specialized for research in multimedia event detection (MED), i.e., automatically identifying what's happening in a video by analyzing the audio and visual content. The videos indexed in the YLI-MED corpus are a subset of the larger YLI feature corpus, which is being developed by the International Computer Science Institute and Lawrence Livermore National Laboratory based on the Yahoo Flickr Creative Commons 100 Million (YFCC100M) dataset. The videos in YLI-MED are categorized as depicting one of ten target events, or no target event, and are annotated for additional attributes like language spoken and whether the video has a musical score. The annotations also include degree of annotator agreement and average annotator confidence scores for the event categorization of each video. Version 1.0 of YLI-MED includes 1823 "positive" videos that depict the target events and 48,138 "negative" videos, as well as 177 supplementary videos that are similar to event videos but are not positive examples. Our goal in producing YLI-MED is to be as open about our data and procedures as possible. This report describes the procedures used to collect the corpus; gives detailed descriptive statistics about the corpus makeup (and how video attributes affected annotators' judgments); discusses possible biases in the corpus introduced by our procedural choices and compares it with the most similar existing dataset, TRECVID MED's HAVIC corpus; and gives an overview of our future plans for expanding the annotation effort.
Motivation & Objective
- To create a large-scale, publicly accessible corpus of user-generated videos annotated for multimedia events to support research in automated event detection.
- To ensure high-quality annotations through multi-stage verification, consensus scoring, and confidence estimation by multiple annotators.
- To enable fair comparison across systems by providing standardized training and test sets with balanced event distribution and negative examples.
- To extend the YFCC100M dataset with rich, structured annotations for events and non-event characteristics such as language and music.
- To lay the foundation for a future Multimedia Genome Project (MMGP) with expanded event coverage and broader annotation of video content.
Proposed method
- Event definitions were developed through iterative research and consensus to ensure clarity and consistency for annotators.
- Videos were collected from YFCC100M using keyword and metadata filters, followed by manual review and annotation for event and non-event characteristics.
- Annotator agreement and confidence scores were computed using inter-annotator agreement metrics and subjective confidence ratings to assess label reliability.
- The corpus was split into training and test sets after removing low-agreement videos and reducing user bias through consensus-based filtering.
- Negative videos were selected via automated metadata filtering and manually verified to ensure they do not depict any target events.
- Computed features—such as audio embeddings (YFCC-CC), visual features (AlexNet), and image features (LIRE)—were pre-computed and released alongside the annotations.
Experimental results
Research questions
- RQ1How can a large-scale, publicly available video corpus be systematically constructed for multimedia event detection with high annotation quality?
- RQ2To what extent do non-event characteristics such as language and music influence event detection performance?
- RQ3How does the distribution of event types and annotator agreement affect the reliability and generalizability of event detection models?
- RQ4What are the key biases introduced by data collection procedures, and how can they be mitigated in corpus design?
- RQ5Can a scalable, collaborative framework for large-scale video annotation (e.g., the Multimedia Genome Project) be effectively launched using this corpus as a foundation?
Key findings
- The YLI-MED v.1.0 corpus contains 2,000 positive videos across 10 target events and 48,700 non-event videos, with a balanced split into training and test sets.
- Annotator agreement for event classification ranged from moderate to substantial, with average confidence scores reflecting high annotator reliability.
- Post-production effects such as music and editing significantly influenced event detection decisions, with music being a notable confounding factor.
- Language characteristics, including non-English speech, were present in a notable fraction of videos and correlated with event classification outcomes.
- The corpus was designed to be comparable to TRECVID MED and HAVIC, though procedural differences may affect direct comparability across datasets.
- Feature bundles for audio, visual, and motion features were pre-computed and released, enabling immediate use in downstream event detection research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.