Skip to main content
QUICK REVIEW

[논문 리뷰] Audeo: Audio Generation for a Silent Performance Video

Kun Su, Xiulong Liu|arXiv (Cornell University)|2020. 06. 23.
Music and Audio Processing참고 문헌 44인용 수 30
한 줄 요약

Audeo는 비디오에서 피아노 상단 시점을 보여주는 퍼포먼스 영상을 예측된 피아노 롤에서 시작해 GAN으로 Pseudo-MIDI로 정제하고, 전통적 또는 딥 MIDI 합성기로 오디오를 합성하는 3단계 전체 파이프라인을 제시한다.

ABSTRACT

We present a novel system that gets as an input video frames of a musician playing the piano and generates the music for that video. Generation of music from visual cues is a challenging problem and it is not clear whether it is an attainable goal at all. Our main aim in this work is to explore the plausibility of such a transformation and to identify cues and components able to carry the association of sounds with visual events. To achieve the transformation we built a full pipeline named ` extit{Audeo}' containing three components. We first translate the video frames of the keyboard and the musician hand movements into raw mechanical musical symbolic representation Piano-Roll (Roll) for each video frame which represents the keys pressed at each time step. We then adapt the Roll to be amenable for audio synthesis by including temporal correlations. This step turns out to be critical for meaningful audio generation. As a last step, we implement Midi synthesizers to generate realistic music. extit{Audeo} converts video to audio smoothly and clearly with only a few setup constraints. We evaluate extit{Audeo} on `in the wild' piano performance videos and obtain that their generated music is of reasonable audio quality and can be successfully recognized with high precision by popular music identification software.

연구 동기 및 목표

  • 피아노 연주에서 보이는 시각적 단서가 대응하는 오디오를 그럴듯하게 생성할 수 있는지 이해한다.
  • 합성에 적합한 음향 표현으로 영상 프레임을 매핑하는 해석 가능한 엔드 투 엔드 파이프라인을 개발한다.
  • 제약이 없는 영상에서 안정적인 오디오 생성을 가능하게 하는 시각적, 기호적, 합성 구성 요소를 식별한다.
  • 실세계 피아노 영상에서 시스템을 평가하고 표준 도구를 사용해 생성된 음악의 인지 가능성을 평가한다.

제안 방법

  • Video2Roll Net은 ResNet18을 기반으로 한 다중 스케일 특징 주의 네트워크로 다섯 연속 흑백 프레임을 처리해 프레임당 눌린 키를 예측하고, 가운데 프레임마다 Piano-Roll M을 출력한다.
  • Roll2Midi Net은 예측된 롤을 GAN(G)과 D로 구성된 네트워크로 정제해 합성에 적합한 강건한 Pseudo-MIDI 표현을 생성한다.
  • Midi Synth는 고전적 MIDI 합성기(FluidSynth)와 딥 합성기(PerfNet 기반 스펙트로그램 정제)를 사용해 MIDI를 스펙트로그램과 Griffin-Lim을 통해 오디오로 변환하며, 선택적 스펙트로그램 정제 단계가 있다.

실험 결과

연구 질문

  • RQ1탑뷰 피아노 영상의 시각적 단서가 오디오 합성을 지원하는 일관된 MIDI 유사 표현으로 번역될 수 있는가?
  • RQ23단계 파이프라인(Video2Roll → Roll2Midi → Midi Synth)이 시각적 연주와 일치하는 들리는 음악을 생성하고 음악 인식 소프트웨어에 의해 감지될 수 있는가?
  • RQ3GAN 기반 정제가 MIDI 예측 품질과 하류 오디오 합성에 미치는 영향은 무엇인가?
  • RQ4예측된 MIDI에서 전통적 합성기와 딥 합성기가 인식 가능하게 정확한 음악을 생성하는 데 어떻게 비교되는가?
  • RQ5제약이 없는 영상과 보지 못한 작곡가에도 시스템이 얼마나 잘 일반화되는가?

주요 결과

  • Video2Roll Net은 ResNet18을 기반으로 한 다중 스케일 특징 주의 네트워크로 다섯 프레임의 연속 흑백 프레임을 처리해 프레임당 눌린 키를 예측하고 가운데 프레임마다 Piano-Roll M을 출력한다.
  • Roll2Midi Net은 예측된 롤을 GAN(G)과 D로 구성된 네트워크로 정제해 합성에 적합한 강건한 Pseudo-MIDI 표현을 생성한다.
  • Midi 합성은 FluidSynth 또는 PerfNet 기반 스펙트로그램 정제를 사용해 MIDI를 스펙트로그램과 Griffin-Lim을 통해 오디오로 변환하며 baselines보다 Midi+FluidSynth와 Midi+PerfNet 구성에서 SoundHound가 더 높은 정확도로 탐지한다.
  • 전체 평가에서 Midi+FluidSynth가 주요 테스트 세트에서 최상의 오디오 탐지율(85.6%)을 달성하며 기준치(ground-truth) 성능(87.7%)에 근접한다.
  • 시스템은 템포, 레가토, 스타카토와 같은 퍼포먼스 변형에 대해 강건함을 보이고 보지 못한 작곡가에게도 일반화되며 비학습 스타일에 대해 탐지율이 약 70%에 가깝다.
  • ResNet 베이스라인과 비교했을 때 전체 Audeo 파이프라인(Video2Roll + Roll2Midi + Midi Synth)은 음악적으로 관련된 지표에서 현저히 우수하다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.