Skip to main content
QUICK REVIEW

[论文解读] Overview of PlantCLEF 2021: cross-domain plant identification

Hervé Goëau, Pierre Bonnet|ArXiv.org|Sep 23, 2025
Genomics and Phylogenetic Studies参考文献 1被引用 50
一句话总结

LifeCLEF 2021 PlantCLEF 挑战评估从标本照片到田野照片的跨域植物识别,强调领域自适应与度量学习以识别罕见物种。

ABSTRACT

Automated plant identification has improved considerably thanks to recent advances in deep learning and the availability of training data with more and more field photos. However, this profusion of data concerns only a few tens of thousands of species, mainly located in North America and Western Europe, much less in the richest regions in terms of biodiversity such as tropical countries. On the other hand, for several centuries, botanists have systematically collected, catalogued and stored plant specimens in herbaria, especially in tropical regions, and recent efforts by the biodiversity informatics community have made it possible to put millions of digitised records online. The LifeCLEF 2021 plant identification challenge (or "PlantCLEF 2021") was designed to assess the extent to which automated identification of flora in data-poor regions can be improved by using herbarium collections. It is based on a dataset of about 1,000 species mainly focused on the Guiana Shield of South America, a region known to have one of the highest plant diversities in the world. The challenge was evaluated as a cross-domain classification task where the training set consisted of several hundred thousand herbarium sheets and a few thousand photos to allow learning a correspondence between the two domains. In addition to the usual metadata (location, date, author, taxonomy), the training data also includes the values of 5 morphological and functional traits for each species. The test set consisted exclusively of photos taken in the field. This article presents the resources and evaluations of the assessment carried out, summarises the approaches and systems used by the participating research groups and provides an analysis of the main results.

研究动机与目标

  • 评估热带植物中标本(源域)与田野照片(目标域)之间跨域植物识别的性能。
  • 评估加入植物性状元数据以改进领域自适应的影响。
  • 提供针对热带地区数据匮乏物种的方法基准与分析。

提出的方法

  • 使用一个大型标本布张数据集(321,270 张标本)与一个较小的田野照片集(6,316 训练;3,186 测试)来学习两个图像域之间的域映射。
  • 引入来自《百科全书生命》(Encyclopedia of Life)的五个功能性性状作为物种层面的元数据来指导训练。
  • 使用全测试集的平均倒数排序(MRR)以及包含少量田野照片的困难子集来评估提交结果。
  • 参与者采用领域自适应和度量学习方法,包括两流式标本-田野影像三元组损失网络(HTFL)和单流混合网络(OSM)。
  • 主办方和参与者用外部数据(如 PlantCLEF2019、GBIF)和多任务线索(属/科性状)来增强模型。
  • 为每个提交提供详细的工作笔记以确保可重复性。

实验结果

研究问题

  • RQ1基于标本的训练能否迁移到田野照片,以用于数据匮乏的热带物种?
  • RQ2基于性状的辅助任务是否提升跨域植物识别性能?
  • RQ3外部训练数据对罕见物种的领域自适应性能有何影响?

主要发现

  • 跨域识别仍然非常具有挑战性;最佳单一模型在全测试集上的 MRR 大约为 0.18–0.20,对困难物种的表现较低。
  • 基于域自适应的 CNN 方法优于传统 CNN,但仍然困难,尤其是对于罕见田野照片物种。
  • 外部数据显著提升性能(例如,当添加外部数据时,主办方的结果从 0.052 提升到 0.153)。
  • 两流式标本-田野三元组损失网络结合集成策略在不同田野照片可用性下实现对多物种的强泛化。
  • 使用属、科和性状信息的多任务扩展在 MR R 上带来可衡量的增益。
  • 主办方的多判别器/子任务提交(性状与分类等级)在主办方提交中实现了最高的初始 MRR。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。