Skip to main content
QUICK REVIEW

[논문 리뷰] Vocal Tract Area Estimation by Gradient Descent

David Südholt, Mateo Cámara|arXiv (Cornell University)|2023. 07. 10.
Speech and Audio ProcessingComputer Science인용 수 3
한 줄 요약

이 논문은 음성의 음향 데이터에서 후각 트랙의 면적 함수와 갈락탈 소스 파arameter를 추정하기 위한 화이트박스, 기울기 기반 최적화 방법을 제안한다. 이는 핑크 트롬본의 운동학적 합성기에서 정확한 음향 재현을 가능하게 한다. 기울기 역전파를 통해 미분 가능한 후각 트랙 모델을 거쳐 오차 기울기를 전파함으로써, 유전 알고리즘과 신경망과 같은 블랙박스 기반 방법보다 주관적 听음 테스트에서 뛰어난 성능을 보였다.

ABSTRACT

Articulatory features can provide interpretable and flexible controls for the synthesis of human vocalizations by allowing the user to directly modify parameters like vocal strain or lip position. To make this manipulation through resynthesis possible, we need to estimate the features that result in a desired vocalization directly from audio recordings. In this work, we propose a white-box optimization technique for estimating glottal source parameters and vocal tract shapes from audio recordings of human vowels. The approach is based on inverse filtering and optimizing the frequency response of a wave\-guide model of the vocal tract with gradient descent, propagating error gradients through the mapping of articulatory features to the vocal tract area function. We apply this method to the task of matching the sound of the Pink Trombone, an interactive articulatory synthesizer, to a given vocalization. We find that our method accurately recovers control functions for audio generated by the Pink Trombone itself. We then compare our technique against evolutionary optimization algorithms and a neural network trained to predict control parameters from audio. A subjective evaluation finds that our approach outperforms these black-box optimization baselines on the task of reproducing human vocalizations.

연구 동기 및 목표

  • 음성 기록물에서 운동학적 파arameter를 추정하기 위한 미분 가능하고 화이트박스 최적화 접근법을 개발하는 것.
  • 핑크 트롬본 운동학적 합성기에서 인간의 음성 표현을 정확하게 재현할 수 있도록 하는 것.
  • 유전 알고리즘과 신경망과 같은 블랙박스 방법의 수렴 속도 및 해석 가능성 측면에서의 한계를 극복하는 것.
  • 기울기 기반 최적화가 시간에 따라 변화하는 음성에 대한 제어 파arameter를 효과적으로 복원할 수 있음을 보여주는 것.
  • 신경망을 사용한 엔드 투 엔드 미분 가능한 운동학적 합성의 기초를 다지는 것.

제안 방법

  • 목표 음성에서 갈락탈 소스 파형과 후각 트랙의 전단 필터 응답을 분리하기 위해 역필터링을 사용한다.
  • 면적 함수에 기반하여 후각 트랙의 전달 함수를 해석적으로 기술하고, 이와 추정된 필터 간의 차이를 최소화하기 위해 경사 하강법을 적용한다.
  • 운동학적 제어 파arameter에서 후각 트랙 면적 함수로의 매핑을 미분 가능하게 구현함으로써 오차 기울기를 파arameter로 역전파할 수 있도록 한다.
  • 최적화는 프레임 단위로 수행되며, 이전 프레임의 해를 초기화로 사용하여 시간적 연속성을 확보한다.
  • 이 방법은 물리 원리를 사용하여 갈락탈 소스와 후각 트랙을 모델링하는 핑크 트롬본 합성기와 적용된다.
  • 비교를 위해 블랙박스 기반 기준으로는 유전 알고리즘, 입자 군집 최적화, 파라미터 예측을 위한 CNN 기반 신경망이 포함된다.
Figure 1: The user interface of the Pink Trombone articulatory synthesizer.
Figure 1: The user interface of the Pink Trombone articulatory synthesizer.

실험 결과

연구 질문

  • RQ1라벨이 없는 훈련 데이터가 없이도 기울기 기반 최적화가 음성에서 후각 트랙 면적 함수를 효과적으로 복원할 수 있는가?
  • RQ2블랙박스 기반 방법과 비교할 때 화이트박스, 미분 가능한 최적화 방법이 인간의 음성 표현을 얼마나 잘 재현하는가?
  • RQ3이 방법은 동적인 음성에 대해 시간에 따라 변화하는 운동학적 파arameter를 정확하게 추정할 수 있는가?
  • RQ4기울기 기반 최적화가 진화적 또는 신경망 기반 접근법보다 더 청각적으로 정확한 재현을 제공하는가?
  • RQ5이 방법을 신경망을 사용한 엔드 투 엔드 미분 가능한 운동학적 합성으로 확장할 수 있는가?

주요 결과

  • 제안된 기울기 기반 방법은 핑크 트롬본 합성기가 자체적으로 생성한 음성에 대해 후각 트랙 면적 함수를 정확히 복원한다.
  • 주관적 听음 테스트에서 이 방법은 모든 블랙박스 기반 기준보다 유의미하게 뛰어난 성능을 보였으며, 유전 알고리즘, 입자 군집 최적화, 훈련된 CNN의 평가 점수를 모두 상회했다 (p < 0.001).
  • 훈련 데이터나 모델의 미세조정 없이도 뛰어난 청각 품질을 달성했다.
  • 프레임 단위 초기화와 기울기 전파를 통해 시간에 따라 변화하는 음성에 대해 매끄럽고 시간적으로 일관된 파arameter 궤적을 확보했다.
  • 프리드먼 검정은 다양한 방법 간의 청각 품질에 유의미한 차이가 있음을 확인했으며 (p < 0.001), 사후 검정은 기울기 기반 접근의 우수성을 확인했다.
  • 이 결과는 운동학적 합성에서 역음향 모델링을 위해 미분 가능한 물리 모델을 활용할 수 있음을 보여준다.
Figure 2: Illustration of the proposed sound matching method. Target audio is inverse filtered to obtain a source waveform and the transfer function of a filter. For resynthesis, the glottal control parameters $F_{0}$ and Tenseness are estimated from the source waveform. The vocal tract area functio
Figure 2: Illustration of the proposed sound matching method. Target audio is inverse filtered to obtain a source waveform and the transfer function of a filter. For resynthesis, the glottal control parameters $F_{0}$ and Tenseness are estimated from the source waveform. The vocal tract area functio

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.