[Paper Review] Overview of ExpertLifeCLEF 2018: how far automated identification systems are from the best experts?
The LifeCLEF 2018 ExpertCLEF compares automated plant-identification systems with top human experts; best AI approaches approach but do not surpass the best experts, with top results around 0.84–0.87 vs experts up to 0.967.
Automated identification of plants and animals has improved considerably in the last few years, in particular thanks to the recent advances in deep learning. The next big question is how far such automated systems are from the human expertise. Indeed, even the best experts are sometimes confused and/or disagree between each others when validating visual or audio observations of living organism. A picture actually contains only a partial information that is usually not sufficient to determine the right species with certainty. Quantifying this uncertainty and comparing it to the performance of automated systems is of high interest for both computer scientists and expert naturalists. The LifeCLEF 2018 ExpertCLEF challenge presented in this paper was designed to allow this comparison between human experts and automated systems. In total, 19 deep-learning systems implemented by 4 different research teams were evaluated with regard to 9 expert botanists of the French flora. The main outcome of this work is that the performance of state-of-the-art deep learning models is now close to the most advanced human expertise. This paper presents more precisely the resources and assessments of the challenge, summarizes the approaches and systems employed by the participating research groups, and provides an analysis of the main outcomes.
Motivation & Objective
- Quantify how close state-of-the-art automated plant identification is to top human experts.
- Create realistic, multi-source training and test datasets including trusted and noisy data.
- Evaluate multiple deep-learning based identification systems on a shared task.
- Analyze cases where machines outperform or underperform human experts.
Proposed method
- Use an ExpertCLEF 2018 task with 19 deep-learning systems from 4 teams.
- Train on trusted (EoL) and noisy web data and test on expert-verified Western European plant observations.
- Evaluate top-1 accuracy for automated runs and compare to expert performance.
- Employ CNN ensembles with data augmentation and test-time averaging.
- Analyze failure cases to understand intrinsic limits of image-based identification.
Experimental results
Research questions
- RQ1How close can deep-learning plant identification get to expert-level accuracy on field-like images?
- RQ2What factors (training data quality, ensembling, data augmentation) most influence machine performance relative to experts?
- RQ3Which observation types or taxonomic groups most distinguish machine from human expert performance?
- RQ4Can automated systems outperform experts on specific, harder cases and why?
Key findings
- Best automated system achieved 0.84 top-1 with expert comparison and 0.867 on the full set.
- Best expert top-1 accuracy ranged from 0.613 to 0.960 with a median of 0.800.
- Automated systems were sometimes better than experts on certain observations (e.g., CMP Run 4 vs best expert on some cases).
- Automated performance approached expert levels but did not surpass the top expert (best expert 0.967 in conclusion).
- Performance gains were linked to training on both trusted and noisy data and using CNN ensembles with data augmentation.
- Several automated runs correctly identified most observations, yet a minority remained challenging due to species similarity or limited information in images.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.