Skip to main content
QUICK REVIEW

[논문 리뷰] Deep Music Analogy Via Latent Representation Disentanglement

Ruihan Yang, Dingsu Wang|arXiv (Cornell University)|2019. 06. 09.
Music and Audio Processing참고 문헌 25인용 수 32
한 줄 요약

논문은 포화된 악상 조건에서 8박 멜로디의 피치와 리듬을 명시적으로 제약된 조건부 변분 자동인코더(EC2-VAE)로 분리하여 잠재 인자를 전이시키고, 객관적 지표와 주관적 연구를 활용해 아날로지를 평가하는 방법을 제시한다.

ABSTRACT

Analogy-making is a key method for computer algorithms to generate both natural and creative music pieces. In general, an analogy is made by partially transferring the music abstractions, i.e., high-level representations and their relationships, from one piece to another; however, this procedure requires disentangling music representations, which usually takes little effort for musicians but is non-trivial for computers. Three sub-problems arise: extracting latent representations from the observation, disentangling the representations so that each part has a unique semantic interpretation, and mapping the latent representations back to actual music. In this paper, we contribute an explicitly-constrained variational autoencoder (EC$^2$-VAE) as a unified solution to all three sub-problems. We focus on disentangling the pitch and rhythm representations of 8-beat music clips conditioned on chords. In producing music analogies, this model helps us to realize the imaginary situation of "what if" a piece is composed using a different pitch contour, rhythm pattern, or chord progression by borrowing the representations from other pieces. Finally, we validate the proposed disentanglement method using objective measurements and evaluate the analogy examples by a subjective study.

연구 동기 및 목표

  • 고수준 추상화를 원시 관찰이 아니라 전이함으로써 아날로지 기반 음악 생성을 가능하게 하려는 동기.
  • 피치와 리듬에 명시적 의미를 갖는 해리된 잠재 공간을 개발한다.
  • 해리가 reconstruction 품질을 크게 저하시키지 않으면서, 추론 시에 유사한 학습 데이터 없이도 아날로지가 가능하도록 한다.

제안 방법

  • 잠재 z를 z_p(피치)와 z_r(리듬)로 해리하기 위한 명시적으로 제약된 조건부 변분 오토인코더(EC2-VAE)를 사용한다.
  • 인코더와 디코더 모두를 코드진행(chords)에 조건화하고, z_r 의미를 강화하기 위한 리듬 중심의 중간 디코더 작업을 포함한다.
  • 잠재 z를 두 부분으로 나누고 z_r을 교차 엔트로피로 학습된 리듬 디코더에 연결하여 리듬 특징과 일치하도록 한다.
  • 해리되어 있는 ELBO 목적이 특정 가정하에 표준 조건부 VAE보다 넉넉하거나 같은 수준으로 유지됨을 보인다.
  • 멜로디를 8비트 시퀀스로 표현하고 피치-온셋 공간은 130차원, 리듬 특징은 3차원으로 표현하며, 코드(chords)는 크롬마 기반 조건으로 제공된다.

실험 결과

연구 질문

  • RQ1VAE 프레임워크 내에서 피치와 리듬을 명시적으로 어떻게 분리할 수 있는가?
  • RQ2분리된 모델이 피치, 리듬, 또는 코드 표현을 서로 다른 곡 사이에서 전이해 의미 있는 아날로지를 가능하게 할 수 있는가?
  • RQ3명시적 해리가 재구성 품질을 저하시킬 수 있는가, 그리고 원래의 ELBO 목표에 근접한 수준으로 남을 수 있는가?
  • RQ4아날로지를 통한 생성의 효과를 보여주는 객관적 및 주관적 증거는 무엇인가?

주요 결과

  • EC2-VAE는 디코더를 피치와 리듬 잠재 요인을 분리하도록 구성함으로써 명시적 해리를 달성한다.
  • 해리가 재구성 품질을 보존하고 코드 조건화와 결합될 수 있어 의미 있는 아날로지를 가능하게 한다.
  • 객관적 지표는 피치와 리듬의 효과적인 분리를 나타내며 (Δz 및 F-score 기반 보강 질의가 의도된 요인과의 높은 정렬을 시사) 고도화된 연관성을 보인다.
  • 질적 예시에서 z_p(피치) 또는 z_r(리듬)을 교체하면서 다른 측면과 코드 조건화를 유지하는 성공적인 아날로지를 보인다.
  • 주관적 평가에서 EC2-VAE 변형은 규칙 기반 기준선보다 더 창의적이고 음악적이지만, 자연스러움과 전반적인 음악성 면에서 여전히 인간이 작곡한 원곡에 뒤처진다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.