[논문 리뷰] Overview of PlantCLEF 2022: Image-based plant identification at global scale
이 논문은 PlantCLEF 2023( LifeCLEF 2023 식물 식별 작업)를 조사하고 CNN과 Vision Transformer를 활용한 대규모 이미지 기반 식물 식별을 분석하며 SSL로 사전학습된 ViT와 학습에 웹 데이터의 이점을 강조합니다. 또한 참가 실행의 주요 결과 표를 제시합니다.
It is estimated that there are more than 300,000 species of vascular plants in the world. Increasing our knowledge of these species is of paramount importance for the development of human civilization (agriculture, construction, pharmacopoeia, etc.), especially in the context of the biodiversity crisis. However, the burden of systematic plant identification by human experts strongly penalizes the aggregation of new data and knowledge. Since then, automatic identification has made considerable progress in recent years as highlighted during all previous editions of PlantCLEF. Deep learning techniques now seem mature enough to address the ultimate but realistic problem of global identification of plant biodiversity in spite of many problems that the data may present (a huge number of classes, very strongly unbalanced classes, partially erroneous identifications, duplications, variable visual quality, diversity of visual contents such as photos or herbarium sheets, etc). The PlantCLEF2022 challenge edition proposes to take a step in this direction by tackling a multi-image (and metadata) classification problem with a very large number of classes (80k plant species). This paper presents the resources and evaluations of the challenge, summarizes the approaches and systems employed by the participating research groups, and provides an analysis of key findings.
연구 동기 및 목표
- 생물다양성 모니터링을 지원하기 위한 자동화된 글로벌 규모의 식물 종 식별의 동기를 제시한다.
- 평가를 위해 테스트 세트와 함께 두 개의 대규모 학습 데이터 세트(신뢰된 데이터와 웹 데이터)를 설명한다.
- 80,000종에 대한 참가 방법, 아키텍처, 학습 전략을 요약한다.
- 향후 대규모 식물 식별 연구를 안내하기 위한 주요 발견을 분석한다.
제안 방법
- 신뢰된 GBIF 유래의 큐레이션 이미지와 노이즈가 있는 웹 이미지를 포함해 총 4백만 장의 이미지로 80k 종에 대해 데이터셋 구성 설명.
- 다중 이미지 관측을 사용한 식물 식별 태스크를 평가하고 매크로 평균된 역순위(MA-MRR)를 메트릭으로 사용.
- 아키텍처 선정(CNN vs Vision Transformers)과 사전학습(STL vs SSL, EVA, MAE)에 중점을 두고 참가 방법 조사.
- 정해진 구성의 참가 실행의 결과를 리더보드에서 보고하고 설정 간 성능 비교.
- 웹 데이터의 영향, 기관 기반 서브모델, 및 사전학습 전략이 MA-MRR에 미치는 영향 분석.
실험 결과
연구 질문
- RQ1대규모 다중 이미지 학습 세트가 식물 종 검색 성능에 미치는 효과는 무엇인가?
- RQ2Vision Transformer가 이 대규모 식물 식별 태스크에서 감독 학습 전이 학습을 사용하는 CNN보다 성능이 우수한가?
- RQ3노이즈가 있는 웹 데이터를 포함하는 것이 다수의 종에 대한 모델 정확도와 일반화에 어떤 영향을 미치는가?
- RQ4대규모 식물 식별에서 최상의 MA-MRR을 얻기 위한 학습 구성(분류 수준, 기관, 사전학습)은 무엇인가?
주요 결과
- Vision Transformer 모델이 자체 감독 학습(EVA/MXE 접근)을 통해 사전학습되었을 때 가장 높은 MA-MRR 점수를 기록하여 CNN 기반 솔루션을 능가했습니다.
- 가장 우수한 EVA 기반 접근(MingleXuRun8)의 MA-MRR은 0.67395였습니다.
- 웹 학습 세트를 도입하면 신뢰 데이터만 사용할 때보다 성능이 크게 향상되었습니다(예: 0.65035에서 0.67395로).
- 학습에서 덜 populutive한 종을 제거하는 종 제거가 성능을 저하시켰으며, 모든 종을 포함하는 것이 중요함을 시사합니다.
- 기관별 모델을 결합하는 것은 일반적으로 종 커버리지가 감소하여 성능을 감소시키는 경향을 보였습니다.
- CNN 기반 방법은 최대 0.61813 MA-MRR에 도달했으나 최고 EVA 결과보다 낮았습니다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.