[论文解读] Overview of ExpertLifeCLEF 2018: how far automated identification systems are from the best experts?
LifeCLEF 2018 ExpertCLEF 将自动植物识别系统与顶尖人类专家进行比较;最佳的 AI 方法接近但未超过最佳专家,顶级结果约为 0.84–0.87,相比之下专家最高达到 0.967。
Automated identification of plants and animals has improved considerably in the last few years, in particular thanks to the recent advances in deep learning. The next big question is how far such automated systems are from the human expertise. Indeed, even the best experts are sometimes confused and/or disagree between each others when validating visual or audio observations of living organism. A picture actually contains only a partial information that is usually not sufficient to determine the right species with certainty. Quantifying this uncertainty and comparing it to the performance of automated systems is of high interest for both computer scientists and expert naturalists. The LifeCLEF 2018 ExpertCLEF challenge presented in this paper was designed to allow this comparison between human experts and automated systems. In total, 19 deep-learning systems implemented by 4 different research teams were evaluated with regard to 9 expert botanists of the French flora. The main outcome of this work is that the performance of state-of-the-art deep learning models is now close to the most advanced human expertise. This paper presents more precisely the resources and assessments of the challenge, summarizes the approaches and systems employed by the participating research groups, and provides an analysis of the main outcomes.
研究动机与目标
- 量化最先进的自动植物识别与顶尖人类专家之间的接近程度。
- 创建现实且多来源的训练与测试数据集,包含可信数据和带噪声的数据。
- 在同一任务上评估多种基于深度学习的识别系统。
- 分析机器在某些情况下胜过或不如人类专家的案例。
提出的方法
- 使用 ExpertCLEF 2018 任务,来自 4 支队伍的 19 个深度学习系统。
- 在可信数据(EoL)和带噪声的网页数据上训练,并在专家验证的西欧植物观测数据上测试。
- 评估自动化运行的 Top-1 准确率,并与专家表现进行比较。
- 采用带数据增强的 CNN 集成和测试时平均。
- 分析失败案例,以理解基于图像的识别的内在极限。
实验结果
研究问题
- RQ1深度学习植物识别在类现场的图像上能接近专家级准确度到什么程度?
- RQ2哪些因素(训练数据质量、集成、数据增强)对机器相对于专家的表现影响最大?
- RQ3哪些观测类型或分类群最能区分机器与人类专家的表现?
- RQ4在特定更难的案例中,自动系统是否能超过专家,原因何在?
主要发现
- 最佳自动化系统在与专家比较时达到 0.84 的 Top-1,在完整集上达到 0.867。
- 最佳专家 Top-1 准确率范围为 0.613 到 0.960,中位数为 0.800。
- 在某些观测中,自动系统有时优于专家(例如在某些案例中 CMP Run 4 相较于最佳专家)。
- 自动化性能已接近专家水平,但未超过顶级专家(结论中最佳专家 0.967)。
- 性能提升与在可信数据和带噪声数据上的训练,以及使用带数据增强的 CNN 集成相关。
- 若干自动化运行正确识别了大多数观测,但少数由于物种相似性或图像信息有限而困难。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。