Skip to main content
QUICK REVIEW

[논문 리뷰] Unsupervised Activity Segmentation by Joint Representation Learning and Online Clustering.

Sateesh Kumar, Sanjay Haresh|arXiv (Cornell University)|2021. 05. 27.
Anomaly Detection Techniques and Applications참고 문헌 90인용 수 5
한 줄 요약

이 논문은 시간적 최적 운반과 시간적 일관성 손실을 사용하여 시퀀스 순서를 유지하고 임bedding 품질을 향상시키는, 비지도 활동 분할을 위한 동시 표현 학습 및 온라인 클러스터링 프레임워크를 제안한다. 비디오 프레임을 미니배치로 온라인으로 처리함으로써 최소한의 메모리 사용으로 최신 기술 수준의 성능을 달성하며, 여러 벤치마크에서 이전 방법들을 능가한다.

ABSTRACT

We present a novel approach for unsupervised activity segmentation, which uses video frame clustering as a pretext task and simultaneously performs representation learning and online clustering. This is in contrast with prior works where representation learning and clustering are often performed sequentially. We leverage temporal information in videos by employing temporal optimal transport and temporal coherence loss. In particular, we incorporate a temporal regularization term into the standard optimal transport module, which preserves the temporal order of the activity, yielding the temporal optimal transport module for computing pseudo-label cluster assignments. Next, the temporal coherence loss encourages neighboring video frames to be mapped to nearby points while distant video frames are mapped to farther away points in the embedding space. The combination of these two components results in effective representations for unsupervised activity segmentation. Furthermore, previous methods require storing learned features for the entire dataset before clustering them in an offline manner, whereas our approach processes one mini-batch at a time in an online manner. Extensive evaluations on three public datasets, i.e. 50-Salads, YouTube Instructions, and Breakfast, and our dataset, i.e., Desktop Assembly, show that our approach performs on par or better than previous methods for unsupervised activity segmentation, despite having significantly less memory constraints.

연구 동기 및 목표

  • 비지도 활동 분할에서 순차적 표현 학습과 클러스터링의 한계를 해결하기 위해.
  • 전체 특징 임베딩을 저장하지 않고도 비디오 프레임을 온라인으로 처리할 수 있도록 하여 메모리 사용을 줄이기 위해.
  • 최적 운반과 일관성 손실을 통해 시간적 구조를 통합함으로써 표현 품질을 향상시키기 위해.
  • 감독 신호나 오프라인 클러스터링에 의존하지 않고도 경쟁 가능한 분할 성능를 달성하기 위해.

제안 방법

  • 프레임 클러스터링을 선제 과제로 사용하여 표현 학습과 온라인 클러스터링을 동시에 수행한다.
  • 가짜 레이블 할당에서 시간 순서 유지 보장을 위해 시간적 최적 운반 모듈을 도입한다.
  • 이웃 프레임은 임베딩 공간에서 가까이, 먼 프레임은 멀리 오게 하기 위해 시간적 일관성 손실을 적용한다.
  • 한 번에 한 미니배치씩 처리함으로써 낮은 메모리 오버헤드로 온라인 학습을 가능하게 한다.
  • 표현 학습과 클러스터링의 공동 최적화를 병합된 손실 구성 요소를 사용해 엔드 투 엔드로 수행한다.

실험 결과

연구 질문

  • RQ1표현 학습과 온라인 클러스터링을 동시에 수행하면 비지도 활동 분할 성능가 향상되는가?
  • RQ2최적 운반을 통해 시간적 구조를 통합하면 클러스터링 품질이 어떻게 향상되는가?
  • RQ3온라인 처리를 통해 얼마나 많은 메모리 사용을 줄일 수 있으며, 성능에 영향을 주지 않는가?
  • RQ4시간적 일관성 손실은 학습된 임베딩의 품질에 어떤 영향을 미치는가?

주요 결과

  • 제안된 방법은 50-Salads, YouTube Instructions, Breakfast 데이터셋에서 이전 최신 기술 수준의 방법들과 비슷하거나 더 뛰어난 성능를 달성한다.
  • 새로 수집한 Desktop Assembly 데이터셋에서도 뛰어난 성능를 보이며, 다양한 도메인으로의 일반화 능력을 입증한다.
  • 오프라인 특징 저장을 피함으로써 메모리 사용을 크게 줄여 온라인 처리를 가능하게 한다.
  • 시간적 최적 운반과 일관성 손실의 통합은 더 시간적으로 일관성 있고 구분력 있는 표현을 이끌어낸다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.