[Paper Review] No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World
The paper analyzes geo-diversity in ImageNet and Open Images, showing amerocentric/eurocentric bias and its impact on classifier performance across regions. It argues for geo-representative data sets for developing-world applications.
Modern machine learning systems such as image classifiers rely heavily on large scale data sets for training. Such data sets are costly to create, thus in practice a small number of freely available, open source data sets are widely used. We suggest that examining the geo-diversity of open data sets is critical before adopting a data set for use cases in the developing world. We analyze two large, publicly available image data sets to assess geo-diversity and find that these data sets appear to exhibit an observable amerocentric and eurocentric representation bias. Further, we analyze classifiers trained on these data sets to assess the impact of these training distributions and find strong differences in the relative performance on images from different locales. These results emphasize the need to ensure geo-representation when constructing data sets for use in the developing world.
Motivation & Objective
- Assess the geo-diversity of two large open image datasets (ImageNet and Open Images).
- Evaluate how training on these datasets affects classifier performance on images from different geographical locations.
- Demonstrate the existence of amerocentric/eurocentric representation biases in widely used datasets.
- Discuss implications for dataset construction in the developing world.
Proposed method
- Use country-level geo-location proxies to estimate geographic distribution in ImageNet and Open Images.
- Analyze the distribution of images across countries and identify representation imbalances.
- Train and evaluate pretrained Inception V3 models on both datasets to compare performance on geographically localized images.
- Collect stress-test data via crowdsourced and geo-located web-image methods to assess classifier behavior across regions.
- Use saliency maps (SmoothGrad) to examine which image regions drive misclassifications.
Experimental results
Research questions
- RQ1Do ImageNet and Open Images exhibit geo-representation biases across countries?
- RQ2How does geographic bias in training data affect classifier performance on non-US images?
- RQ3Are misclassifications influenced more by attire or context in region-specific images?
- RQ4Do classifiers trained on these datasets perform consistently across different geographic locales?
Key findings
- Open Images and ImageNet show substantial US- and Europe-skew, with China and India underrepresented.
- A large portion of samples come from the six most represented countries in North America and Europe.
- Classifiers trained on these datasets misclassify region-specific images more often, with lower confidence for non-US images.
- Saliency maps indicate that models rely on facial regions rather than attire for some misclassifications.
- Stress-tested images from Hyderabad often have lower likelihoods under both models, indicating regional performance gaps.
- Differences in performance across countries suggest non-uniform representation across image categories.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.