Skip to main content
QUICK REVIEW

[논문 리뷰] Delving Deep into the Generalization of Vision Transformers under Distribution Shifts

Chongzhi Zhang, Mingyuan Zhang|arXiv (Cornell University)|2021. 06. 14.
Domain Adaptation and Few-Shot Learning참고 문헌 31인용 수 9
한 줄 요약

이 논문은 다양한 분포 이탈(Out-of-Distribution, OOD) 상황에서 비전 트랜스포머(Vision Transformers, ViTs)의 일반화 성능을 조사하며, ViTs가 인간 인지와 유사한 형태와 구조에 대한 더 강한 인도적 편향(inductive biases)을 학습함으로써 기존의 CNNs보다 더 뛰어난 OOD 성능을 보임을 밝혀냈다. 적대적 훈련, 정보 이론, 그리고 매끄러운 훈련을 통한 자기지도 학습을 활용해 일반화 향상된 ViTs(GE-ViTs)를 도입함으로써, 기존의 ViTs 대비 OOD 데이터에서 4%의 정확도 향상을 달성했으며, 크기가 큰 모델일수록 이점이 더 두드러졌다.

ABSTRACT

Vision Transformers (ViTs) have achieved impressive performance on various vision tasks, yet their generalization under distribution shifts (DS) is rarely understood. In this work, we comprehensively study the out-of-distribution (OOD) generalization of ViTs. For systematic investigation, we first present a taxonomy of DS. We then perform extensive evaluations of ViT variants under different DS and compare their generalization with Convolutional Neural Network (CNN) models. Important observations are obtained: 1) ViTs learn weaker biases on backgrounds and textures, while they are equipped with stronger inductive biases towards shapes and structures, which is more consistent with human cognitive traits. Therefore, ViTs generalize better than CNNs under DS. With the same or less amount of parameters, ViTs are ahead of corresponding CNNs by more than 5% in top-1 accuracy under most types of DS. 2) As the model scale increases, ViTs strengthen these biases and thus gradually narrow the in-distribution and OOD performance gap. To further improve the generalization of ViTs, we design the Generalization-Enhanced ViTs (GE-ViTs) from the perspectives of adversarial learning, information theory, and self-supervised learning. By comprehensively investigating these GE-ViTs and comparing with their corresponding CNN models, we observe: 1) For the enhanced model, larger ViTs still benefit more for the OOD generalization. 2) GE-ViTs are more sensitive to the hyper-parameters than their corresponding CNN models. We design a smoother learning strategy to achieve a stable training process and obtain performance improvements on OOD data by 4% from vanilla ViTs. We hope our comprehensive study could shed light on the design of more generalizable learning architectures.

연구 동기 및 목표

  • 실세계의 분포 이탈 상황에서 비전 트랜스포머(ViTs)의 OOD 일반화 행동을 이해하는 것.
  • 다양한 종류의 분포 이탈 유형에서 ViT의 일반화 성능를 CNNs와 비교하는 것.
  • ViTs가 학습하는 인도적 편향을 규명하고, 이들이 인간의 인지적 특성(예: 형태 및 구조에 대한 민감도)과 어떻게 일치하는지 밝혀내는 것.
  • 적대적 학습, 정보 이론, 자기지도 학습을 활용해 OOD 일반화를 향상시키는 Generalization-Enhanced ViTs(GE-ViTs)를 설계하고 평가하는 것.
  • GE-ViTs의 하이퍼파rameter 민감도를 분석하고 안정적인 수렴을 위해 더 매끄러운 훈련 전략을 제안하는 것.

제안 방법

  • 의미적 개념 수정 기반의 분포 이탈 유형 분류 체계를 제안: 배경, 손상, 질감, 스타일 이탈.
  • 다섯 가지 OOD 이탈 유형에서 ViT 변종과 대응하는 CNNs(예: DeiT, VGG-16, BiT)에 대한 광범위한 평가 수행.
  • 일반화 성능 향상을 위한 세 가지 유형의 GE-ViTs 도입: T-ADV(적대적 훈련), T-MME(상호정보량 최대화), T-SSL(자기지도 학습).
  • 특히 적대적 및 정보이론 기반 방법에서 중요한 안정성 확보를 위해 매끄러운 훈련 전략 적용.
  • 여러 벤치마크에서 통일된 평가 프로토콜을 사용해 OOD 성능 및 일반화 갭을 비교.
  • 표준 평가 지표(예: OOD 이탈 벤치마크에서의 상위-1 정확도)를 활용해 성능 향상을 검증.

실험 결과

연구 질문

  • RQ1ViTs는 다양한 종류의 분포 이탈 상황에서 CNNs와 비교해 어떻게 일반화되는가?
  • RQ2ViTs는 어떤 인도적 편향을 학습하며, 이는 인간의 인지적 특성(예: 형태 및 구조에 대한 민감도)과 어떻게 일치하는가?
  • RQ3모델 스케일링이 ViTs의 OOD 일반화 성능에 어느 정도 기여하는가?
  • RQ4일반화 향상된 ViTs(GE-ViTs)는 다양한 OOD 이탈 유형에서 일관된 성능 향상을 달성할 수 있는가?
  • RQ5GE-ViTs는 하이퍼파rameter에 얼마나 민감한가? 더 매끄러운 훈련 전략은 이들의 훈련 수렴을 안정화시킬 수 있는가?

주요 결과

  • ViTs는 대부분의 분포 이탈 상황에서 CNNs보다 더 뛰어난 일반화 성능를 보이며, 유사하거나 더 적은 파라미터로도 OOD 데이터에서 5% 이상 높은 상위-1 정확도를 달성한다.
  • ViTs는 형태와 구조에 대한 더 강한 인도적 편향을 보이며, 질감과 배경에 대한 편향은 약간으로, 인간의 시각 인지와 더 유사하게 작동한다.
  • 모델 스케일링이 증가할수록 ViTs의 형태/구조 편향이 강화되어, 내부 분포와 외부 분포 간의 성능 격차가 점차 좁아진다.
  • GE-ViTs는 기존의 ViTs 대비 OOD 데이터에서 4%의 성능 향상을 기록했으며, 세 가지 향상 방법 모두에서 일관된 개선 효과를 보였다.
  • 크기가 큰 ViTs일수록 일반화 향상 기법에서 더 큰 이점을 얻었으며, 이는 OOD 일반화 성능에서의 확장성 우수성을 시사한다.
  • GE-ViTs는 CNN 대비 하이퍼파rameter에 더 민감하며, 안정적인 수렴을 위해 더 매끄러운 훈련 전략이 필요하다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.