[Paper Review] Meta-Album: Multi-domain Meta-Dataset for Few-Shot Image Classification
Meta-Album introduces a large-scale, multi-domain meta-dataset for few-shot image classification, comprising 40 diverse datasets across 10 domains—such as ecology, manufacturing, and optical character recognition—each with at least 20 classes and 40 images per class. It offers three scalable versions (Micro, Mini, Extended) with uniform formatting, verified licenses, and extensibility via community contributions, enabling realistic benchmarking of meta-learning and transfer learning algorithms across heterogeneous domains.
We introduce Meta-Album, an image classification meta-dataset designed to facilitate few-shot learning, transfer learning, meta-learning, among other tasks. It includes 40 open datasets, each having at least 20 classes with 40 examples per class, with verified licences. They stem from diverse domains, such as ecology (fauna and flora), manufacturing (textures, vehicles), human actions, and optical character recognition, featuring various image scales (microscopic, human scales, remote sensing). All datasets are preprocessed, annotated, and formatted uniformly, and come in 3 versions (Micro $\subset$ Mini $\subset$ Extended) to match users' computational resources. We showcase the utility of the first 30 datasets on few-shot learning problems. The other 10 will be released shortly after. Meta-Album is already more diverse and larger (in number of datasets) than similar efforts, and we are committed to keep enlarging it via a series of competitions. As competitions terminate, their test data are released, thus creating a rolling benchmark, available through OpenML.org. Our website https://meta-album.github.io/ contains the source code of challenge winning methods, baseline methods, data loaders, and instructions for contributing either new datasets or algorithms to our expandable meta-dataset.
Motivation & Objective
- To address the lack of diverse, realistic, and computationally feasible meta-datasets for few-shot learning.
- To enable benchmarking of meta-learning algorithms in cross-domain settings that reflect real-world distribution shifts.
- To create a scalable, extensible meta-dataset that supports lightweight, medium, and large-scale experiments.
- To facilitate community-driven expansion through competitions and open contributions of new datasets and algorithms.
- To provide a rolling benchmark via released competition test sets on OpenML.org, ensuring long-term utility and reproducibility.
Proposed method
- The meta-dataset integrates 40 preprocessed, annotated, and uniformly formatted image classification datasets from 10 distinct domains, including fauna, flora, textures, vehicles, and OCR.
- Datasets are curated with at least 20 classes and 40 images per class, ensuring sufficient data for few-shot learning while maintaining diversity in scale and modality.
- Three versions (Micro, Mini, Extended) are provided to accommodate varying computational resources, with increasing data volume and coverage.
- All datasets are licensed for academic use, with verified open licenses and metadata available through the project website and OpenML.org.
- A community contribution pipeline is established, including code for data formatting, validation, and a review process to ensure quality and consistency.
- The meta-dataset is integrated into a competition framework (NeurIPS 2021 and 2022), with test sets released post-competition to form a rolling benchmark on OpenML.
Experimental results
Research questions
- RQ1Can a multi-domain meta-dataset improve the generalization evaluation of few-shot learning algorithms beyond single-domain benchmarks?
- RQ2How does the inclusion of diverse domains and class hierarchies affect the performance and robustness of meta-learning models?
- RQ3To what extent can a community-driven, extensible meta-dataset sustain long-term benchmarking utility in meta-learning research?
- RQ4Does the availability of uniform, preprocessed data across domains reduce the barrier to entry for new researchers in few-shot learning?
- RQ5How do different data scales (Micro, Mini, Extended) impact the performance and training efficiency of meta-learning models?
Key findings
- Meta-Album includes 40 image classification datasets across 10 domains, with 30 available immediately and 10 released in spring 2023, offering a broader and more diverse collection than prior meta-datasets.
- The Micro version contains 32,000 images across 40 datasets with 19–20 classes per domain and 40 images per class, totaling 380 MB, making it lightweight and accessible.
- The Mini version scales to 220,950 images across the same 40 datasets, with 19–706 classes per domain and 40 images per class, totaling 3.9 GB, suitable for most research systems.
- The Extended version contains 1.58 million images, with class counts ranging from 19 to 706 and image counts per class from 1 to 187, totaling 15 GB, supporting large-scale experiments.
- Meta-Album is the only meta-dataset designed for extensibility, with open-source tools and a review process to support ongoing community contributions of new datasets and algorithms.
- Test data from NeurIPS 2021 and 2022 competitions are released via OpenML.org, creating a rolling benchmark that evolves over time and supports reproducible research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.