[論文レビュー] Audeo: Audio Generation for a Silent Performance Video
Audeoは、ビデオからPiano-Rollを予測し、それをGANでPseudo-MIDIへ refin した後、伝統的または深層MIDI合成で音声を合成する、トップビューのピアノ演奏ビデオを合成音楽に変換する3段階のパイプラインを提示します。
We present a novel system that gets as an input video frames of a musician playing the piano and generates the music for that video. Generation of music from visual cues is a challenging problem and it is not clear whether it is an attainable goal at all. Our main aim in this work is to explore the plausibility of such a transformation and to identify cues and components able to carry the association of sounds with visual events. To achieve the transformation we built a full pipeline named ` extit{Audeo}' containing three components. We first translate the video frames of the keyboard and the musician hand movements into raw mechanical musical symbolic representation Piano-Roll (Roll) for each video frame which represents the keys pressed at each time step. We then adapt the Roll to be amenable for audio synthesis by including temporal correlations. This step turns out to be critical for meaningful audio generation. As a last step, we implement Midi synthesizers to generate realistic music. extit{Audeo} converts video to audio smoothly and clearly with only a few setup constraints. We evaluate extit{Audeo} on `in the wild' piano performance videos and obtain that their generated music is of reasonable audio quality and can be successfully recognized with high precision by popular music identification software.
研究の動機と目的
- ピアノ演奏の視覚的手がかりから対応する音声を仮説的に生成できるかを理解する。
- 映像フレームを合成に適した音楽表現へマッピングする、解釈可能なエンドツーエンドのパイプラインを開発する。
- 制約のない動画から頑健な音声生成を可能にする視覚的・記号的・合成要素を特定する。
- 実世界のピアノ動画でシステムを評価し、標準ツールを用いて生成音楽の認識性を評価する。
提案手法
- Video2Roll Netは五帧のグレースケール画像を処理し、ResNet18を基盤としたマルチスケール特徴アテンションネットワークで中間フレームごとにPiano-Roll Mを出力して、フレームごとの押下キーを予測する。
- Roll2Midi NetはGAN(生成器G、識別器D)を用いて予測されたRollを refinementし、合成に適した頑健なPseudo-MIDI表現を生成する。
- Midi Synthは古典的なMIDI合成(FluidSynth)と深層合成子(PerfNetベースのスペクトログラム refin)を用いて、スペクトログラムとGriffin-Limによる音声変換を行い、任意のスペクトログラム refinステップを提供する。
実験結果
リサーチクエスチョン
- RQ1トップビューのピアノ動画から視覚的手がかりを、音声合成をサポートする一貫したMIDI様表現へ翻訳できるか。
- RQ23段階パイプライン(Video2Roll → Roll2Midi → Midi Synth)は、視覚演奏と一致する聴覚音楽を生成し、音楽識別ソフトウェアで検出可能か。
- RQ3GANベースの refin が MIDI予測品質と下流の音声合成に与える影響はどの程度か。
- RQ4伝統的および深層合成器は、予測されたMIDIから認識可能に正確な音楽を生み出す点でどのように比較されるか。
- RQ5制約のない動画と未知の作曲家への一般化性能はどの程度か。
主な発見
- Video2Roll Netは、フレームレベルのキー予測において従来手法と比較してリコールが高く、精度も競争力があり、MIDIの正確性を向上させた。
- Roll2Midi Netは偽陽性・偽陰性を削減し、F1スコアを向上させ、MIDI予測精度を高めた。
- FluidSynthまたはPerfNetベースのスペクトログラム refinを用いたMIDI合成は、SoundHoundがMidi+FluidSynthおよび Midi+PerfNet構成を基準より高い精度で検出した。
- 全体的な評価で、Midi+FluidSynthはメインのテストセットで最良の音声検出率(85.6%)を達成し、グラウンドトゥルース性能(87.7%)に近づいた。
- 演奏の変動(テンポ、レガート、スタッカート)に対する頑健性を示し、未知の作曲家にも一般化し、非トレーニングスタイルの検出率は約70%。
- ResNetベースラインと比較して、完全なAudeoパイプライン(Video2Roll + Roll2Midi + Midi Synth)は、音楽的に関連する指標で大幅に上回った。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。