Skip to main content
QUICK REVIEW

[論文レビュー] AI Choreographer: Music Conditioned 3D Dance Generation with AIST++

Ruilong Li, Shan Yang|arXiv (Cornell University)|Jan 21, 2021
Human Motion and Animation参考文献 81被引用数 32
ひとこと要約

AIST++(多モーダルな大規模3Dダンスモーションデータセット)と、長い音楽条件付き3Dダンスシーケンスを2秒のシードモーションから生成するFull-Attention Cross-modal TransformerであるFACTを紹介。

ABSTRACT

We present AIST++, a new multi-modal dataset of 3D dance motion and music, along with FACT, a Full-Attention Cross-modal Transformer network for generating 3D dance motion conditioned on music. The proposed AIST++ dataset contains 5.2 hours of 3D dance motion in 1408 sequences, covering 10 dance genres with multi-view videos with known camera poses -- the largest dataset of this kind to our knowledge. We show that naively applying sequence models such as transformers to this dataset for the task of music conditioned 3D motion generation does not produce satisfactory 3D motion that is well correlated with the input music. We overcome these shortcomings by introducing key changes in its architecture design and supervision: FACT model involves a deep cross-modal transformer block with full-attention that is trained to predict $N$ future motions. We empirically show that these changes are key factors in generating long sequences of realistic dance motion that are well-attuned to the input music. We conduct extensive experiments on AIST++ with user studies, where our method outperforms recent state-of-the-art methods both qualitatively and quantitatively.

研究の動機と目的

  • 現実的な音楽条件付き3Dダンス生成の必要性を動機づけ、このタスクを研究するための大規模データセットを提供する。
  • シードシーケンスが短くても入力音楽と整列した長い3Dダンスモーションを生成できるモデルを開発する。
  • 未来N superviseを用いた全アテンションのクロスモーダル融合が安定した高品質のモーション生成をもたらすことを示す。
  • SMPLベースの表現を通じて生成モーションを新規キャラクターへ容易にリターゲット可能にする。

提案手法

  • AIST++を提案する:3Dダンスモーション5.2時間、1408シーケンス、10ジャンルのデータセットで、カメラ姿勢が既知のマルチビュー動画から取得。
  • FACTを導入する:オーディオ変換器、シードモーション変換器、全アテンションを備えたクロスモーダル変換器の3つの要素から成る深いクロスモーダル変換器。早期融合により両方のモダリティを統合。
  • FACTを自動回帰設定でN未来フレーム(N=20、実験)を予測するよう訓練し、future-N supervisionを用いて長距離生成を安定化。
  • 3Dダンスを関節回転とグローバル平移として表現し、新規キャラクターへのリターゲットをサポート。
  • オーディオとモーションの埋め込みの早期融合と12層のクロスモーダル変換器を用いて、音楽とモーションの相関を効果的に学習。
Figure 2: Cross-Modal Music Conditioned 3D Motion Generation Overview. Our proposed a Full-Attention Cross-modal Transformer (FACT) network (details in Figure 3 ) takes in a music piece and a $2$ -second sequence of seed motion, then auto-regressively generates long-range future motions that correla
Figure 2: Cross-Modal Music Conditioned 3D Motion Generation Overview. Our proposed a Full-Attention Cross-modal Transformer (FACT) network (details in Figure 3 ) takes in a music piece and a $2$ -second sequence of seed motion, then auto-regressively generates long-range future motions that correla

実験結果

リサーチクエスチョン

  • RQ1AIST++のような大規模なマルチビュー データセットは、音楽条件付き3Dダンス生成の堅牢な学習を可能にするか。
  • RQ2全アテンションのクロスモーダル変換器とfuture-N supervisionは、因果アテンションや浅い融合のベースラインより、長く音楽と整合したダンスシーケンスの生成で優れているか。
  • RQ3音楽とモーションモダリティの早期融合は、モデルが音楽の手がかりに従う能力にどのように影響するか。
  • RQ4生成されたシーケンスのモーション品質、多様性、音楽-ダンスの整合性を捉える指標はどれが最適か。

主な発見

  • AIST++は、30名の被験者、1408シーケンス、10ジャンル、同期したマルチビュー画像を備える大規模でマルチビュー・マルチジャンルのデータセットとして検証されている。
  • 全アテンション、future-N supervision、早期クロスモーダル融合を備えるFACTは、ベースラインより長く、フリーズせず、音楽と整合した3Dダンスモーションを生成する。
  • FACTは、AIST++上で、Li et al. 2021、Dancenet、DanceRevolutionと比較して、モーションの現実性(FID_kおよびFID_gの低下)とBeatAlignスコアが優れている。
  • ユーザ研究では、FACT生成モーションがベースラインより音楽の一貫性が高く知覚される割合が、Li et al.との比較で81%、Dancenetで71%、DanceRevolutionで77%である。
  • アブレーション研究は、質の高い音楽条件付き生成のために未来N supervision付き全アテンションと早期クロスモーダル融合が不可欠であることを確認した。
Figure 3: FACT Model Details. (a) The structure of the audio/motion/cross-modal transformer with $N$ attention layers. (b) Attention and supervision mechanism as a simplified two-layer model. Models like GPT [ 66 ] and the motion generator of [ 55 ] use causal attention (left) to predict the immedia
Figure 3: FACT Model Details. (a) The structure of the audio/motion/cross-modal transformer with $N$ attention layers. (b) Attention and supervision mechanism as a simplified two-layer model. Models like GPT [ 66 ] and the motion generator of [ 55 ] use causal attention (left) to predict the immedia

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。