Skip to main content
QUICK REVIEW

[논문 리뷰] LogoDet-3K: A Large-Scale Image Dataset for Logo Detection

Jing Wang, Weiqing Min|arXiv (Cornell University)|2020. 08. 12.
Advanced Image and Video Retrieval Techniques참고 문헌 42인용 수 9
한 줄 요약

이 논문은 3,000개의 카테고리, 194,261개의 주석이 부여된 객체, 158,652张의 이미지를 포함한 가장 큰 공개 로고 검출 데이터셋인 LogoDet-3K를 소개한다. Focal Loss와 CIoU Loss를 통합한 YOLOv3 기반의 Logo-Yolo 방법을 제안하여 클래스 불균형과 바운딩 박스 회귀 문제를 해결하였으며, LogoDet-3K에서 55.37% mAP를 달성하여 YOLOv3보다 4% 높은 성능을 보이며 다른 데이터셋에서도 강력한 일반화 능력을 입증하였다.

ABSTRACT

Logo detection has been gaining considerable attention because of its wide range of applications in the multimedia field, such as copyright infringement detection, brand visibility monitoring, and product brand management on social media. In this paper, we introduce LogoDet-3K, the largest logo detection dataset with full annotation, which has 3,000 logo categories, about 200,000 manually annotated logo objects and 158,652 images. LogoDet-3K creates a more challenging benchmark for logo detection, for its higher comprehensive coverage and wider variety in both logo categories and annotated objects compared with existing datasets. We describe the collection and annotation process of our dataset, analyze its scale and diversity in comparison to other datasets for logo detection. We further propose a strong baseline method Logo-Yolo, which incorporates Focal loss and CIoU loss into the state-of-the-art YOLOv3 framework for large-scale logo detection. Logo-Yolo can solve the problems of multi-scale objects, logo sample imbalance and inconsistent bounding-box regression. It obtains about 4% improvement on the average performance compared with YOLOv3, and greater improvements compared with reported several deep detection models on LogoDet-3K. The evaluations on other three existing datasets further verify the effectiveness of our method, and demonstrate better generalization ability of LogoDet-3K on logo detection and retrieval tasks. The LogoDet-3K dataset is used to promote large-scale logo-related research and it can be found at https://github.com/Wangjing1551/LogoDet-3K-Dataset.

연구 동기 및 목표

  • 높은 다양성과 복잡성을 지닌 대규모로 완전 주석이 부여된 로고 검출 데이터셋의 부족 문제를 해결하기 위해.
  • 카테고리 커버리지와 객체 주석 규모를 증가시켜 로고 검출에 더 도전적인 벤치마크를 구축하기 위해.
  • 다중 척도, 불균형, 복잡한 배경 등의 과제에 대응하기 위해 대규모 로고 검출에 특화된 강력한 베이스라인 모델인 Logo-Yolo를 개발하기 위해.
  • 다양한 데이터셋과 작업(검출 및 검색)을 통해 LogoDet-3K의 일반화 능력을 검증하기 위해.

제안 방법

  • LogoDet-3K는 로고 이미지 수집, 필터링, 158,652장의 이미지에 걸쳐 194,261개의 바운딩 박스에 대한 수동 주석을 포함한 철저한 파이프라인을 통해 구축된다.
  • 이 데이터셋은 로고 유형, 변환, 배경 복잡성 측면에서 높은 다양성을 지닌 3,000개의 고유한 로고 카테고리를 포함한다.
  • Logo-Yolo는 YOLOv3 기반으로 구축되며, 로고 검출에서 긴 꼬리 클래스 불균형 문제를 완화하기 위해 Focal Loss를 통합한다.
  • CIoU Loss는 작은 로고나 왜곡된 로고에 특히 유리한 바운딩 박스 회귀 정확도 향상을 위해 Logo-Yolo에 통합된다.
  • 이 방법은 LogoDet-3K에서 학습 및 평가되며, 일반화 능력을 평가하기 위해 기존의 세 가지 데이터셋(QMUL-OpenLogo, FlickrLogos-32, WebLogo-2M)에서도 테스트된다.
  • 제거 분석을 통해 Focal Loss와 CIoU Loss가 도전적인 로고 인스턴스의 검출 성능 향상에 효과적임을 확인한다.
Figure 1: Statistics of LogoDet-3K categories and images. The abscissa represents the number of logo images, the ordinate represents the number of categories.
Figure 1: Statistics of LogoDet-3K categories and images. The abscissa represents the number of logo images, the ordinate represents the number of categories.

실험 결과

연구 질문

  • RQ1기존의 로고 검출 데이터셋과 비교해 볼 때, LogoDet-3K는 규모, 다양성, 주석 품질 측면에서 어떻게 다른가?
  • RQ2수정된 YOLOv3 프레임워크는 작은 로고, 다중 척도, 불균형 로고 검출 과제를 효과적으로 처리할 수 있는가?
  • RQ3제안된 Logo-Yolo 모델은 새로운 벤치마크와 기존의 벤치마크 데이터셋 모두에서 뛰어난 성능과 일반화 능력을 달성하는가?
  • RQ4LogoDet-3K에서의 사전 훈련이 로고 검색 작업의 성능 향상에 얼마나 기여하는가?

주요 결과

  • LogoDet-3K는 3,000개의 로고 카테고리, 194,261개의 주석이 부여된 객체, 158,652장의 이미지를 포함하여 지금까지 가장 큰 완전 주석 로고 검출 데이터셋이다.
  • Logo-Yolo는 LogoDet-3K에서 55.37% mAP를 달성하여 YOLOv3보다 4% 높은 성능을 보이며, 동일한 벤치마크에서 다른 최신 기술 모델들을 능가한다.
  • 이 방법은 강력한 일반화 능력을 보이며, QMUL-OpenLogo에서 LogoDet-3K로 미세조정했을 때 56.10% mAP를 기록하여 데이터셋 간 이식성이 뛰어나다.
  • ResNet101을 사용하여 FlickrLogos-32에서의 검색 성능을 향상시키기 위해 LogoDet-3K에서 사전 훈련을 수행한 결과, mAP가 52.62%에서 54.17%로 상승하였다.
  • 실패 사례 분석을 통해 매우 작은, 가림을 입거나 매우 유사한 로고를 탐지하는 데 어려움이 있음을 확인하였으며, 이는 다중 레이블 및 소형 객체 검출 문제에 여전히 도전 과제가 있음을 시사한다.
  • Focal Loss와 CIoU Loss의 통합은 특히 희귀 및 소형 로고 인스턴스의 검출 정확도 향상에 크게 기여한다.
Figure 2: Image samples from various categories of LogoDet-3K.
Figure 2: Image samples from various categories of LogoDet-3K.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.