Skip to main content
QUICK REVIEW

[논문 리뷰] Overview of PlantCLEF 2021: cross-domain plant identification

Hervé Goëau, Pierre Bonnet|ArXiv.org|2025. 09. 23.
Genomics and Phylogenetic Studies참고 문헌 1인용 수 50
한 줄 요약

LifeCLEF 2021 PlantCLEF 대회는 열대 식물의 표본집에서 현장 사진까지의 교차 도메인 식물 식별을 평가하며, 도메인 적응 및 메트릭 학습에 중점을 두고 희귀 종 식별을 강조한다.

ABSTRACT

Automated plant identification has improved considerably thanks to recent advances in deep learning and the availability of training data with more and more field photos. However, this profusion of data concerns only a few tens of thousands of species, mainly located in North America and Western Europe, much less in the richest regions in terms of biodiversity such as tropical countries. On the other hand, for several centuries, botanists have systematically collected, catalogued and stored plant specimens in herbaria, especially in tropical regions, and recent efforts by the biodiversity informatics community have made it possible to put millions of digitised records online. The LifeCLEF 2021 plant identification challenge (or "PlantCLEF 2021") was designed to assess the extent to which automated identification of flora in data-poor regions can be improved by using herbarium collections. It is based on a dataset of about 1,000 species mainly focused on the Guiana Shield of South America, a region known to have one of the highest plant diversities in the world. The challenge was evaluated as a cross-domain classification task where the training set consisted of several hundred thousand herbarium sheets and a few thousand photos to allow learning a correspondence between the two domains. In addition to the usual metadata (location, date, author, taxonomy), the training data also includes the values of 5 morphological and functional traits for each species. The test set consisted exclusively of photos taken in the field. This article presents the resources and evaluations of the assessment carried out, summarises the approaches and systems used by the participating research groups and provides an analysis of the main results.

연구 동기 및 목표

  • 표본집(소스)과 현장 사진(타깃) 간의 열대 식물의 교차 도메인 식물 식별 성능을 평가한다.
  • 도메인 적응을 개선하기 위해 식물 형질 메타데이터를 추가하는 영향을 평가한다.
  • 데이터가 부족한 열대 지역 종에 대한 방법과 결과의 벤치마크 및 분석을 제공한다.

제안 방법

  • 큰 표본집 표본 데이터 세트(321,270개 표본)와 더 작은 현장 사진 세트(학습 6,316; 테스트 3,186)를 사용해 두 이미지 도메인 간의 도메인 매핑을 학습한다.
  • Encyclopedia of Life에서 species 수준 메타데이터로 다섯 가지 기능 형질을 도입해 훈련을 안내한다.
  • 전체 테스트 세트와 현장 사진이 적은 난이도 하위 집합에서 Mean Reciprocal Rank(MRR)로 제출을 평가한다.
  • 참가자들은 도메인 적응 및 메트릭 학습 접근법을 사용했으며, 두-스트림 표본-사진 트리플렛 로스 네트워크(HTFL)와 한 스트림 혼합 네트워크(OSM)를 포함한다.
  • 주최측과 참가자들은 외부 데이터(예: PlantCLEF2019, GBIF)와 다중 작업 신호(속(genus)/가족 트레이츠)를 사용해 모델을 보강했다.
  • 재현성을 보장하기 위한 제출별 상세 작업 노트를 제공한다.

실험 결과

연구 질문

  • RQ1표본집 기반의 학습이 데이터가 부족한 열대 종의 현장 사진으로 전이될 수 있는가?
  • RQ2형질 기반 보조 작업이 교차 도메인 식물 식별 성능을 개선하는가?
  • RQ3외부 학습 데이터가 희귀 종의 도메인 적응 성능에 미치는 영향은 무엇인가?

주요 결과

  • 교차 도메인 식별은 여전히 매우 도전적이다; 최상의 단일 모델은 전체 테스트 세트에서 약 0.18–0.20의 MRR에 도달하며, 어려운 종에서 성능은 더 낮다.
  • 도메인 적응 CNN기반 접근법이 전통적인 CNN보다 성능이 우수하지만 여전히 어려움을 겪으며, 특히 희귀한 현장 사진 종에서 문제를 보인다.
  • 외부 데이터는 성능을 크게 향상시킨다(예: 외부 데이터 추가 시 주최자 실행이 0.052에서 0.153로 상승).
  • 두-스트림 표본-사진 트리플렛 로스 네트워크에 앙상블 전략이 현장 사진 이용 가능성이 달라진 종 전반에 걸쳐 강한 일반화성을 보인다.
  • 다중 작업 확장(속, 가족, 형질 정보)이 MRR에서 측정 가능한 이점을 제공한다.
  • 주최자의 다-감별기/하위 작업 제출(형질 및 분류 계통 수준)이 주최자 실행 중 최상위 MRR을 얻었다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.