[논문 리뷰] Controllable cross-speaker emotion transfer for end-to-end speech synthesis.
이 논문은 두 개의 감정 분리 모듈을 통해 감정과 발화자 신원을 분리함으로써, 타코트론2 기반의 엔드 투 엔드 음성 합성 프레임워크를 제안한다. 감정 강도를 제어할 수 있는 학습 가능한 스칼라를 도입하여, 감정 유출가 줄어든 Prosody가 다양한 감정 표현이 가능한 음성을 타겟 발화자에게 제공함으로써 최신 기술 수준의 성능을 달성한다.
The cross-speaker emotion transfer task in TTS particularly aims to synthesize speech for a target speaker with the emotion transferred from reference speech recorded by another (source) speaker. During the emotion transfer process, the identity information of the source speaker could also affect the synthesized results, resulting in the issue of speaker leakage. This paper proposes a new method with the aim to synthesize controllable emotional expressive speech and meanwhile maintain the target speaker's identity in the cross-speaker emotion TTS task. The proposed method is a Tacotron2-based framework with the emotion embedding as the conditioning variable to provide emotion information. Two emotion disentangling modules are contained in our method to 1) get speaker-independent and emotion-discriminative embedding, and 2) explicitly constrain the emotion and speaker identity of synthetic speech to be that as expected. Moreover, we present an intuitive method to control the emotional strength in the synthetic speech for the target speaker. Specifically, the learned emotion embedding is adjusted with a flexible scalar value, which allows controlling the emotion strength conveyed by the embedding. Extensive experiments have been conducted on a Mandarin disjoint corpus, and the results demonstrate that the proposed method is able to synthesize reasonable emotional speech for the target speaker. Compared to the state-of-the-art reference embedding learned methods, our method gets the best performance on the cross-speaker emotion transfer task, indicating that our method achieves the new state-of-the-art performance on learning the speaker-independent emotion embedding. Furthermore, the strength ranking test and pitch trajectories plots demonstrate that the proposed method can effectively control the emotion strength, leading to prosody-diverse synthetic speech.
연구 동기 및 목표
- 엔드 투 엔드 TTS에서 감정과 발화자 신원을 분리함으로써 타발화자 감정 전이 시 발생하는 발화자 유출 문제를 해결한다.
- 타겟 발화자의 합성 음성에서 감정 표현의 강도를 정밀하게 제어할 수 있도록 한다.
- 발화자 독립적인 감정 임베딩을 학습하여 감정 음성 합성의 품질과 Prosody 다양성을 향상시킨다.
- 타겟 발화자 신원을 유지하면서도 타발화자 감정 전이에서 최신 기술 수준의 성능을 달성한다.
제안 방법
- 프레임워크는 타코트론2 기반이며, 감정 임베딩을 조건 변수로 사용하여 감정 표현을 유도하는 음성 합성에 활용된다.
- 두 개의 감정 분리 모듈을 도입한다: 하나는 발화자 독립적이고 감정 구별 가능한 임베딩을 추출하고, 다른 하나는 출력의 감정과 발화자 신원을 명시적으로 제약한다.
- 감정 임베딩에 학습 가능한 스칼라를 적용하여 합성 음성에서 감정 표현의 강도를 제어한다.
- 생성된 음성의 발화자 임베딩이 타겟 발화자와 일치하도록 명시적으로 제약함으로써 신원 일관성을 확보한다.
- 감정 표현력과 발화자 신원 유지 최적화를 위해 중국어 분리 코퍼스를 기반으로 엔드 투 엔드로 학습한다.
- 스칼라 제어 메커니즘을 통해 감정 강도를 조절함으로써 Prosody가 다양한 출력을 가능하게 한다.
실험 결과
연구 질문
- RQ1소스 발화자로부터 타겟 발화자로 감정을 효과적으로 전이하면서도 발화자 유출을 최소화할 수 있는가?
- RQ2감정 표현의 강도를 분리되고 직관적인 방식으로 합성 음성에서 제어할 수 있는가?
- RQ3발화자 독립적인 감정 임베딩을 효과적으로 학습할 수 있는가?
- RQ4제안된 방법이 기존의 최신 기술 기반 참조 임베딩 접근법보다 타발화자 감정 전이 성능에서 뛰어나게 되는가?
주요 결과
- 기존의 참조 임베딩 기반 방법과 비교해 본 논문의 방법은 타발화자 감정 전이 작업에서 최신 기술 수준의 성능을 달성한다.
- 강도 순위 테스트 결과, 스칼라 제어 메커니즘이 합성 문장 간 감정 강도를 효과적으로 조절함을 확인하였다.
- 피치 경로도는 다양한 감정 표현을 가진 Prosody가 다양한 음성을 생성함을 보여준다.
- 두 개의 감정 분리 모듈은 감정과 발화자 신원을 분리함으로써 발화자 유출를 성공적으로 감소시켰다.
- 모델는 타겟 발화자에게 자연스럽고 타당한 감정 표현 음성을 합성하며, 높은 발화자 신원 유지 정확도를 유지한다.
- 조절 가능한 감정 강도를 가진 감정 표현 음성 생성에서 이 방법은 강건성과 제어 가능성을 보였다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.