Skip to main content
QUICK REVIEW

[논문 리뷰] Florence: A New Foundation Model for Computer Vision

Lu Yuan, Dongdong Chen|arXiv (Cornell University)|2021. 11. 22.
Multimodal Machine Learning Applications참고 문헌 55인용 수 340
한 줄 요약

Florence는 장면에서 객체로, 이미지에서 비디오로, RGB에서 다중 모달리티로 표현을 확장하는 대규모 시각-언어 기반 파운데이션 모델로, 최첨단 전이 성능과 광범위한 태스크 적응성을 달성합니다.

ABSTRACT

Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human vision. Computer vision foundation models, which are trained on diverse, large-scale dataset and can be adapted to a wide range of downstream tasks, are critical for this mission to solve real-world computer vision applications. While existing vision foundation models such as CLIP, ALIGN, and Wu Dao 2.0 focus mainly on mapping images and textual representations to a cross-modal shared representation, we introduce a new computer vision foundation model, Florence, to expand the representations from coarse (scene) to fine (object), from static (images) to dynamic (videos), and from RGB to multiple modalities (caption, depth). By incorporating universal visual-language representations from Web-scale image-text data, our Florence model can be easily adapted for various computer vision tasks, such as classification, retrieval, object detection, VQA, image caption, video retrieval and action recognition. Moreover, Florence demonstrates outstanding performance in many types of transfer learning: fully sampled fine-tuning, linear probing, few-shot transfer and zero-shot transfer for novel images and objects. All of these properties are critical for our vision foundation model to serve general purpose vision tasks. Florence achieves new state-of-the-art results in majority of 44 representative benchmarks, e.g., ImageNet-1K zero-shot classification with top-1 accuracy of 83.74 and the top-5 accuracy of 97.18, 62.4 mAP on COCO fine tuning, 80.36 on VQA, and 87.8 on Kinetics-600.

연구 동기 및 목표

  • 공간-시간-모달리티 축에 걸친 다양한 태스크를 위한 어댑터를 더한 사전 학습 모델로서의 컴퓨터 비전 파운데이션 모델 정의.
  • 두 타워 아키텍처를 갖춘 통합된 웹-스케일 이미지-텍스트 사전 학습 프레임워크 구축.
  • 객체 수준, 비디오 및 비전-언어 태스크용 어댑터를 개발하여 광범위한 전이 가능성을 실현.
  • 대규모 데이터셋에서 사전 학습을 효율적으로 확장하기 위한 학습 인프라 최적화.

제안 방법

  • 필터링과 UniCL 기반의 통합 이미지-텍스트 대비 학습으로 9억 개의 이미지-텍스트 쌍 데이터셋(FLD-900M)을 큐레이션.
  • 이미지 인코더(CoSwin/Hierarchical ViT)와 언어 인코더(12-layer transformer)를 사용하는 두-tower Florence 모델을 UniCL을 이용하여 이미지-라벨-설명 공간에서 사전 학습.
  • Dynamic Head 어댑터와 FLOD-9M을 통해 객체 수준 표현 확장 및 객체 탐지 사전 학습.
  • 섬세한 융합을 위한 METER 어댑터를 사용하여 V+L 능력을 통합하고 ITM 및 MLM 손실로 사전 학습.
  • 2D를 3D 토큰으로 변환하고 어텐션/포지셔널 임베딩을 조정하여 Video CoSwin 어댑터로 비디오에 적응.
  • 대용량 배치 및 대규모 학습을 가능하게 하는 확장 가능한 학습 기법(ZeRO, 활성화 체크포인팅, 혼합 정밀도, 그래디언트 캐시)을 시연.

실험 결과

연구 질문

  • RQ1공간-시간-모달리티에 걸친 진정한 컴퓨터 비전 파운데이션 모델은 무엇으로 구성되는가?
  • RQ2하나의 사전 학습 모델이 경량 어댑터를 통해 제로샷, 소수샷, 전체 파인튜닝 방식에서 분류, 검색, 탐지, VQA, 캡션 생성, 비디오 태스크를 포함한 다양한 CV 태스크에서 최첨단 성능을 달성할 수 있는가?
  • RQ3웹-스케일 이미지-텍스트 데이터와 통합 학습 목표가 비전 태스크 및 모달리티 간 전이 가능성에 어떤 영향을 미치는가?

주요 결과

  • Florence는 ImageNet-1K 제로샷 상위 1위 83.74 및 상위 5위 97.18를 포함한 44개 대표 벤치마크에서 새로운 최첨단 성능을 달성합니다.
  • COCO 미세조정에서 62.4 mAP를 달성; VQA 점수는 80.36에 도달; Kinetics-600은 87.8% 정확도를 달성합니다.
  • 제로샷 전이가 평가 세트의 12개 분류 태스크 중 9개에서 우수하고, 선형 탐색은 11개 데이터세트 중 9개에서 우승합니다.
  • Flickr30K 및 MSCOCO에서 제로샷 이미지-텍스트 검색은 경쟁적에서 우수한 결과를 낳으며, Florence가 이전 제로샷 방법들을 능가합니다.
  • FLOD-9M 및 Dynamic Head를 이용한 객체 탐지에서 COCO 및 기타 탐지 벤치마크에서 강력한 AP를 달성합니다(예: COCO 미세조정에서 AP 62.0).
  • Florence는 CD-FSL 벤치마크에서 강력한 도메인 간 소수샷 성능을 보이며, 여러 설정에서 이전 단일 모델 베이스라인을 능가합니다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.