[논문 리뷰] Unsupervised Object Representation Learning using Translation and Rotation Group Equivariant VAE
TARGET-VAE는 인코더 및 생성망에서 군 등변성 조건을 부여함으로써 회전 및 이동에 대해 불변인 물체 표현을 학습하는 완전히 비지도 학습 변동형 오토인코더이다. 이는 의미, 회전, 이동 잠재변수를 동시에 추론하여, 심지어 심한 손상된 이미지에서도 비지도 자세 추정 및 의미 클러스터링 분야에서 최고 성능을 달성한다.
In many imaging modalities, objects of interest can occur in a variety of locations and poses (i.e. are subject to translations and rotations in 2d or 3d), but the location and pose of an object does not change its semantics (i.e. the object's essence). That is, the specific location and rotation of an airplane in satellite imagery, or the 3d rotation of a chair in a natural image, or the rotation of a particle in a cryo-electron micrograph, do not change the intrinsic nature of those objects. Here, we consider the problem of learning semantic representations of objects that are invariant to pose and location in a fully unsupervised manner. We address shortcomings in previous approaches to this problem by introducing TARGET-VAE, a translation and rotation group-equivariant variational autoencoder framework. TARGET-VAE combines three core innovations: 1) a rotation and translation group-equivariant encoder architecture, 2) a structurally disentangled distribution over latent rotation, translation, and a rotation-translation-invariant semantic object representation, which are jointly inferred by the approximate inference network, and 3) a spatially equivariant generator network. In comprehensive experiments, we show that TARGET-VAE learns disentangled representations without supervision that significantly improve upon, and avoid the pathologies of, previous methods. When trained on images highly corrupted by rotation and translation, the semantic representations learned by TARGET-VAE are similar to those learned on consistently posed objects, dramatically improving clustering in the semantic latent space. Furthermore, TARGET-VAE is able to perform remarkably accurate unsupervised pose and location inference. We expect methods like TARGET-VAE will underpin future approaches for unsupervised object generation, pose prediction, and object detection.
연구 동기 및 목표
- 감독 없이 이미지 내 물체의 이동 및 회전에 대해 불변인 의미적 물체 표현을 학습하는 것.
- 기본 VAE가 의미 콘텐츠에서 자세 및 위치를 분리하는 데에 한계가 있음을 해결하는 것.
- 공간 변환에 대한 구조적 인도적 편향을 통합하여 비지도 자세 및 위치 추론을 향상시키는 것.
- 객체가 무작위로 회전 및 위치가 설정된 영상 모odalities(예: 크로이-전자현미경)에서도 강건한 표현 학습을 가능하게 하는 것.
- 완전히 미분 가능하고 종단 간 엔드 투 엔드 프레임워크를 개발하여 분리성과 재구성 성능를 동시에 최적화하는 것.
제안 방법
- 변환 특성에 따라 특징을 추출하기 위해 회전 및 이동에 대해 군 등변성인 인코더 네트워크를 활용한다.
- 잠재변수를 의미, 회전, 이동 구성요소로 분리하는 구조적으로 분리된 변동형 사후 분포를 사용한다.
- 분리된 잠재변수에서 이미지를 재구성하기 위해 공간적으로 등변성인 생성망 네트워크를 구현한다.
- 회전에 대해 균일한 사전을 적용하고, 모든 잠재변수의 사후 분포를 근사하기 위해 공동 추론 네트워크를 사용한다.
- 인코더 및 생성망에서 가중치 공유 및 군 컨볼루션 연산을 통해 등변성을 강제한다.
- 자세나 위치에 대한 감독 없이도 관측된 이미지만을 사용하여 전체 모델을 종단 간 엔드 투 엔드로 훈련한다.
실험 결과
연구 질문
- RQ1완전히 비지도 VAE는 물체의 이동 및 회전에 대해 불변인 분리된 표현을 학습할 수 있는가?
- RQ2감독 없이도 모델은 물체의 자세(회전 및 이동)를 정확히 추론할 수 있는가?
- RQ3분리된 표현은 의미적 물체 클래스의 후행 클러스터링 성능를 향상시키는가?
- RQ4실세계 영상 데이터(예: 크로이-전자현미경 현미경도)와 같은 고노이즈 및 높은 변동성 환경에서도 이 프레임워크는 일반화 가능한가?
- RQ5군 등변성 설계는 기존 표준 VAE에 비해 표현 품질을 어떻게 향상시키는가?
주요 결과
- TARGET-VAE는 훈련 이미지가 심하게 회전 및 이동으로 인해 손상된 상황에서도 일관된 자세를 가진 데이터에서 학습한 표현과 거의 동일한 정확도로 의미 표현을 학습한다.
- 모델은 감독 없이도 고품질의 비지도 자세 추정을 달성하여 물체의 회전 및 이동을 정확히 추론한다.
- 잠재공간 내 의미 클러스터링은 알려진 의미 레이블과 일치하며, 크로이-전자현미경 데이터셋에서 입자 시야 및 오염물질이 명확히 분리된다.
- EMPIAR-10025에서 TARGET-VAE는 서로 다른 입자 시야 및 오염 상태에 해당하는 세 개의 명확한 클러스터를 식별한다.
- EMPIAR-10029에서는 입자 시야의 연속적인 변화를 탐지하고, 회전 및 이동에 관계없이 노이즈 제거된 재구성 이미지를 생성한다.
- 이 프레임워크는 이전 방법들에 비해 분리성에서 크게 뛰어나며, 자세-의미 엔트레인먼트와 같은 일반적인 병태를 피한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.