Skip to main content
QUICK REVIEW

[Paper Review] Rehabilitating the ColorChecker Dataset for Illuminant Estimation

Ghalia Hemrit, Graham D. Finlayson|arXiv (Cornell University)|May 30, 2018
Color Science and Applications3 citations
TL;DR

This paper identifies critical flaws in the widely used ColorChecker dataset for illuminant estimation, including multiple inconsistent ground-truths and calculation errors. The authors propose a new, corrected 'recommended' ground-truth (REC) using the Shi and Funt methodology with verified code, demonstrating that algorithm rankings based on legacy ground-truths are unreliable and often reversed under the new standard, thereby rehabilitating the dataset for future benchmarking.

ABSTRACT

In a previous work, it was shown that there is a curious problem with the benchmark ColorChecker dataset for illuminant estimation. To wit, this dataset has at least 3 different sets of ground-truths. Typically, for a single algorithm a single ground-truth is used. But then different algorithms, whose performance is measured with respect to different ground-truths, are compared against each other and then ranked. This makes no sense. We show in this paper that there are also errors in how each ground-truth set was calculated. As a result, all performance rankings based on the ColorChecker dataset - and there are scores of these - are inaccurate. In this paper, we re-generate a new 'recommended' set of ground-truth based on the calculation methodology described by Shi and Funt. We then review the performance evaluation of a range of illuminant estimation algorithms. Compared with the legacy ground-truths, we find that the difference in how algorithms perform can be large, with many local rankings of algorithms being reversed. Finally, we draw the readers attention to our new 'open' data repository which, we hope, will allow the ColorChecker set to be rehabilitated and once again to become a useful benchmark for illuminant estimation algorithms.

Motivation & Objective

  • To identify and correct inconsistencies in the ground-truth illuminant values used in the benchmark ColorChecker dataset for illuminant estimation.
  • To resolve the problem of multiple, conflicting ground-truth sets (SFU, Gt1, Gt2) that lead to unreliable algorithm performance rankings.
  • To establish a definitive, transparent, and reproducible ground-truth for the ColorChecker dataset using the methodology of Shi and Funt, corrected for prior calculation errors.
  • To re-evaluate the performance of 23 illuminant estimation algorithms using the new recommended ground-truth (REC) and compare results with legacy ground-truths.
  • To promote the adoption of the new REC ground-truth as the standard benchmark for future illuminant estimation research.

Proposed method

  • Re-implement the Shi and Funt ground-truth calculation methodology using custom, open-source code to ensure transparency and reproducibility.
  • Identify and correct specific errors in prior ground-truth calculations, including incorrect bounding box detection and inconsistent selection of achromatic patches for white point estimation.
  • Generate a new recommended ground-truth (REC) set for the 568 images in the ColorChecker dataset based on the corrected application of the Shi and Funt method.
  • Compare the REC ground-truth with existing legacy ground-truths (SFU, Gt1, Gt2) using chromaticity distribution plots and statistical analysis.
  • Re-evaluate 23 illuminant estimation algorithms using the REC ground-truth and compare performance rankings with those based on SFU and Gt1.
  • Host the new REC ground-truth and associated code in an open, community-accessible repository to enable ongoing validation and improvement.

Experimental results

Research questions

  • RQ1Why do performance rankings of illuminant estimation algorithms vary significantly when evaluated on different ground-truths from the ColorChecker dataset?
  • RQ2What specific errors in the calculation of legacy ground-truths lead to inconsistent and unreliable algorithm rankings?
  • RQ3How does the proposed recommended ground-truth (REC) compare to existing ground-truths in terms of chromaticity distribution and consistency?
  • RQ4To what extent do algorithm rankings change when using the new REC ground-truth instead of legacy ground-truths?
  • RQ5Can the proposed REC ground-truth serve as a stable, reliable, and open benchmark for future illuminant estimation research?

Key findings

  • The legacy ground-truths (SFU, Gt1, Gt2) are inconsistent and contain calculation errors, including incorrect bounding boxes and non-uniform white point selection across RGB channels.
  • The new recommended ground-truth (REC) aligns closely with SFU but corrects specific outliers, resulting in a more accurate and consistent dataset.
  • Performance rankings of illuminant estimation algorithms are drastically different when using Gt1 compared to REC, with the top 5 algorithms under Gt1 being among the worst under REC.
  • Algorithms such as Fast Fourier Color Constancy and Deep Color Constancy using CNNs are ranked 1st and 6th under REC but are 15th and top-ranked under Gt1, respectively, indicating that training data bias affects performance evaluation.
  • The median angular error rankings for REC and SFU are nearly identical, confirming that REC is a reliable correction of SFU, but Gt1 produces significantly different and misleading results.
  • The study demonstrates that prior performance evaluations using multiple ground-truths are fundamentally flawed, and the REC ground-truth provides the first consistent and definitive benchmark for the ColorChecker dataset.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.