[Paper Review] KonIQ-10k: Towards an ecologically valid and large-scale IQA database
This paper presents KonIQ-10k, a large-scale, ecologically valid IQA database with 10,073 images and 1.2 million crowd-rated quality scores, designed to improve blind image quality assessment in-the-wild.
The main challenge in applying state-of-the-art deep learning methods to predict image quality in-the-wild is the relatively small size of existing quality scored datasets. The reason for the lack of larger datasets is the massive resources required in generating diverse and publishable content. We present a new systematic and scalable approach to create large-scale, authentic and diverse image datasets for Image Quality Assessment (IQA). We show how we built an IQA database, KonIQ-10k, consisting of 10,073 images, on which we performed very large scale crowdsourcing experiments in order to obtain reliable quality ratings from 1,467 crowd workers (1.2 million ratings). We argue for its ecological validity by analyzing the diversity of the dataset, by comparing it to state-of-the-art IQA databases, and by checking the reliability of our user studies.
Motivation & Objective
- Create a large, authentic IQA database representative of real-world internet photos.
- Ensure content and distortion diversity through scalable sampling from a large image collection.
- Obtain reliable subjective quality scores via crowdsourcing and validate their reliability against expert ratings.
- Demonstrate the database’s value by benchmarking existing NR-IQA methods and highlighting the impact of dataset size on model performance.
Proposed method
- Sampling 10,073 images from ~4.8M YFCC100m entries using seven quality indicators and one content indicator plus machine tags.
- Uniform sampling across indicators via MILP-based optimization and a 200-bin discretization per indicator.
- Deep-feature-based content sampling using 4096-dimensional VGG-16 FC7 features for diversity.
- Crowdsourcing with 1,467 workers to obtain 120 ratings per image, plus expert-based validation.
- Quality control via test questions, screening for agreement with global MOS, and removal of low-quality or duplicate data.
- Evaluation of no-reference IQA methods on KonIQ-10k, LIVE In the Wild, and TID2013 using SROCC and PLCC.
Experimental results
Research questions
- RQ1How to construct a large-scale, ecologically valid IQA database that captures real-world content and distortions?
- RQ2What sampling strategies can ensure content and distortion diversity at scale?
- RQ3Are crowdsourced MOS reliable for high-quality subjective scores when properly screened?
- RQ4Do existing NR-IQA methods perform differently on ecologically valid, natural-image datasets compared to artificial-distortion databases?
- RQ5How does dataset size impact the performance of blind IQA models?
Key findings
- KonIQ-10k contains 10,073 images with 1.2 million crowd ratings from 1,467 workers.
- Crowd MOS aligns with expert ratings with an RMSE of 11.35 on a 100-point scale, and 73% of images are within expert-consistent error bounds.
- The proposed sampling yields greater diversity in brightness, colorfulness, contrast, sharpness, and content distribution than prior databases.
- NR-IQA methods show stronger correlations on KonIQ-10k and LIVE In the Wild than on TID2013, highlighting dataset size and natural content impact on performance.
- Larger training data improves IQA model performance, as evidenced by better NR-IQA predictions on KonIQ-10k.
- The table of NR-IQA method performance shows KonIQ-10k achieving higher SROCC/PLCC than some baselines on the in-the-wild datasets.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.