[논문 리뷰] CALM: Contrastive Aligned Audio-Language Multirate and Multimodal Representations
CALM은 사전 훈련된 언어 모델 임bedding의 공유 임bedding 공간에 음성 스펙트로그램 패치를 정렬하기 위해 스펙트럼 트랜스포머와 공동 음성-언어 사전 훈련을 통해 대비적이고 다속도이며 다모odal인 프레임워크를 제안한다. 이는 단 몇 시간의 훈련과 8개의 V100 GPU를 사용함에도 불구하고 자동 음성 인식(transcripts)이 있음에도 불구하고 기존 시스템 대비 10–25% 향상된 성능을 기록하며 최신 기술 수준(SOTA)을 달성한다.
Deriving multimodal representations of audio and lexical inputs is a central problem in Natural Language Understanding (NLU). In this paper, we present Contrastive Aligned Audio-Language Multirate and Multimodal Representations (CALM), an approach for learning multimodal representations using contrastive and multirate information inherent in audio and lexical inputs. The proposed model aligns acoustic and lexical information in the input embedding space of a pretrained language-only contextual embedding model. By aligning audio representations to pretrained language representations and utilizing contrastive information between acoustic inputs, CALM is able to bootstrap audio embedding competitive with existing audio representation models in only a few hours of training time. Operationally, audio spectrograms are processed using linearized patches through a Spectral Transformer (SpecTran) which is trained using a Contrastive Audio-Language Pretraining objective to align audio and language from similar queries. Subsequently, the derived acoustic and lexical tokens representations are input into a multimodal transformer to incorporate utterance level context and derive the proposed CALM representations. We show that these pretrained embeddings can subsequently be used in multimodal supervised tasks and demonstrate the benefits of the proposed pretraining steps in terms of the alignment of the two embedding spaces and the multirate nature of the pretraining. Our system shows 10-25\% improvement over existing emotion recognition systems including state-of-the-art three-modality systems under various evaluation objectives.
연구 동기 및 목표
- 사전 훈련된 언어 모델의 공유 임bedding 공간에 맞춰진 음성-언어 표현을 학습하는 것.
- 말하기 언어 입력으로부터 다속도 및 대비적 정보를 활용하여 표현 학습의 효율성과 성능을 향상시키는 것.
- 저자원 또는 노이즈가 많은 번역 조건에서도 최소한의 피지테이닝으로 종단 간 다모달 이해를 가능하게 하는 것.
- 음성 표현이 언어 모델의 지도를 통해 효과적으로 부트스트랩될 수 있으며, 대규모 음성 전용 사전 훈련에 대한 의존도를 줄일 수 있음을 보여주는 것.
제안 방법
- 선형화된 스펙트로그램 패치를 처리하여 음성 토큰을 생성하는 스펙트럼 트랜스포머(SpecTran)로 짧은 시간의 음성 프레임에 대한 패치 기반 자기주의 어텐션을 가능하게 한다.
- 대비적 음성-언어 사전 훈련(CALP)은 유사한 쿼리 쌍에 대한 대비 목표를 사용하여 음성 임베딩을 해당 사전 훈련된 어휘 임베딩과 정렬한다.
- 다모달 트랜스포머는 음성 및 언어 토큰을 통합하여 문장 수준의 맥락을 포함하여 공동 CALM 표현을 생성한다.
- 음성 및 언어 모odal의 공동 사전 훈련을 위해 마스크된 언어 모델링(MLM)과 마스크된 음성 모델링(MAM) 손실을 결합한다.
- 모델은 단일 모달 추론(음성 전용 또는 언어 전용)을 지원하며, 하류 작업에서 최소한의 피지테이닝으로 종단 간 엔드 투 엔드로 훈련된다.
실험 결과
연구 질문
- RQ1사전 훈련된 언어 모델의 임베딩 공간에 맞춰 음성 표현을 효과적으로 학습할 수 있는가?
- RQ2단기 및 장기 다속도 정보와 대비적 정보를 활용하면 음성-언어 표현 학습이 향상되는가?
- RQ3경량의 대비적 사전 훈련 접근 방식이 최소한의 훈련 시간으로도 정서 인식에서 최신 기술 수준의 성능을 달성할 수 있는가?
- RQ4높은 단어 오류율을 가진 자동 음성 인식(transcripts)을 사용할 경우 이 접근 방식의 강건성은 어떠한가?
주요 결과
- CALM은 다양한 평가 목표에서 기존의 정서 인식 시스템, 특히 최신 기술 수준의 3모달 모델들 대비 10–25% 상대적 향상을 기록한다.
- CMU-MOSEI 및 UTD MSP-Podcasts 데이터셋 모두에서 베이스라인을 초월하며, 사전 훈련만으로도 CMU-MOSEI에서 가중 정확도 2%p의 절대적 향상을 기록한다.
- CALM의 BERT_TINY 버전조차도 다른 최신 기술 수준 알고리즘을 능가하는 성능을 보이며, 파rameter 수가 감소한 상태에서도 강력한 성능을 입증한다.
- CALM 사전 훈련은 8개의 V100 GPU에서 3시간 이내에 완료되어 다른 다모달 접근 방식에 비해 높은 계산 효율성을 보여준다.
- 자동 음성 인식(transcripts)으로 생성된 번역이 노이즈가 많더라도 성능이 뛰어나, 이는 모델이 높은 강건성을 지닌다는 것을 입증한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.