[Paper Review] The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
Open Images V4 provides a unified, large-scale dataset with 9.2M images, 30.1M image-level labels for 19.8k concepts, 15.4M bounding boxes for 600 object classes, and 375k visual relationship annotations across 57 relation classes.
We present Open Images V4, a dataset of 9.2M images with unified annotations for image classification, object detection and visual relationship detection. The images have a Creative Commons Attribution license that allows to share and adapt the material, and they have been collected from Flickr without a predefined list of class names or tags, leading to natural class statistics and avoiding an initial design bias. Open Images V4 offers large scale across several dimensions: 30.1M image-level labels for 19.8k concepts, 15.4M bounding boxes for 600 object classes, and 375k visual relationship annotations involving 57 classes. For object detection in particular, we provide 15x more bounding boxes than the next largest datasets (15.4M boxes on 1.9M images). The images often show complex scenes with several objects (8 annotated objects per image on average). We annotated visual relationships between them, which support visual relationship detection, an emerging task that requires structured reasoning. We provide in-depth comprehensive statistics about the dataset, we validate the quality of the annotations, we study how the performance of several modern models evolves with increasing amounts of training data, and we demonstrate two applications made possible by having unified annotations of multiple types coexisting in the same images. We hope that the scale, quality, and variety of Open Images V4 will foster further research and innovation even beyond the areas of image classification, object detection, and visual relationship detection.
Motivation & Objective
- Offer a large-scale, CC-BY licensed dataset collected from Flickr with no preselected class list to reduce bias and enable cross-task research.
- Provide unified annotations for image classification, object detection, and visual relationship detection within the same images.
- Deliver extensive statistical analyses, annotation quality validation, and baseline explorations of model performance as training data scales.
- Demonstrate applications made possible by unified annotations, including fine-grained detection and zero-shot visual relationship detection.
Proposed method
- Collect ~9.2M Flickr images with CC-BY license and filter for privacy/bias reduction including removal of duplicates and non-web-wide images.
- Define 19,794 image-level concepts and 600 boxable object classes (with a hierarchical structure) for annotations.
- Annotate image-level labels via a computer-assisted workflow combining multiple image classifiers and human verification.
- Annotate 15.4M bounding boxes for 600 object classes using extreme-clicking and box-verification series, including hierarchical deduplication and attribute tagging.
- Annotate 374.8k visual relationship triplets by selecting object pairs that plausibly realize relationships and verifying them, including non-trivial, non-co-occurrence-based relations.
- Provide a data-collection and annotation pipeline suitable for cross-task training and analysis across classification, detection, and visual relationships.
Experimental results
Research questions
- RQ1How large-scale, unified annotations across classification, detection, and visual relationship tasks can be collected and validated in a single dataset?
- RQ2What are the statistics, quality characteristics, and biases of Open Images V4 compared to prior datasets?
- RQ3How do modern models' performances evolve with increasing amounts of training data on this scale?
- RQ4What new cross-task applications become feasible with unified annotations (e.g., fine-grained detection without explicit box labels, zero-shot relationship detection)?
Key findings
- Open Images V4 contains 9.18M images, 30.11M image-level labels for 19,794 concepts, 15.44M bounding boxes for 600 object classes, and 374.77k visual relationship triplets across 57 relation classes.
- On average, images contain 8 annotated objects, and bounding boxes total more than 15× the size of the next largest datasets (15.4M boxes on 1.9M images).
- The dataset emphasizes complex scenes and CC-BY licensing to enable broad use, including commercial contexts, while enabling cross-task research with unified annotations.
- Quality validation analyzes geometric box accuracy and annotation recall, and model baselines illustrate performance trends as data scale increases.
- Two novel applications enabled by unified annotations are demonstrated: fine-grained object detection without fine-grained box labels and zero-shot visual relationship detection.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.