[論文レビュー] HSPACE: Synthetic Parametric Humans Animated in Complex Environments
HSPACEは、複雑な屋内・屋外環境に配置された多様でパrametricに変化する人間の、大規模かつ写真のようにリアルな合成データセットを提供する。このデータセットは、GHUMボディモデルを用いて、リアルな動きと完全な3Dの真値を備えたアニメーションを実現している。このデータセットを弱い教師付きの実データと組み合わせることで、モデルの容量が増加するほど、3D人間のポーズとボディシェイプ推定の性能が顕著に向上する。
Advances in the state of the art for 3d human sensing are currently limited by the lack of visual datasets with 3d ground truth, including multiple people, in motion, operating in real-world environments, with complex illumination or occlusion, and potentially observed by a moving camera. Sophisticated scene understanding would require estimating human pose and shape as well as gestures, towards representations that ultimately combine useful metric and behavioral signals with free-viewpoint photo-realistic visualisation capabilities. To sustain progress, we build a large-scale photo-realistic dataset, Human-SPACE (HSPACE), of animated humans placed in complex synthetic indoor and outdoor environments. We combine a hundred diverse individuals of varying ages, gender, proportions, and ethnicity, with hundreds of motions and scenes, as well as parametric variations in body shape (for a total of 1,600 different humans), in order to generate an initial dataset of over 1 million frames. Human animations are obtained by fitting an expressive human body model, GHUM, to single scans of people, followed by novel re-targeting and positioning procedures that support the realistic animation of dressed humans, statistical variation of body proportions, and jointly consistent scene placement of multiple moving people. Assets are generated automatically, at scale, and are compatible with existing real time rendering and game engines. The dataset with evaluation server will be made available for research. Our large-scale analysis of the impact of synthetic data, in connection with real data and weak supervision, underlines the considerable potential for continuing quality improvements and limiting the sim-to-real gap, in this practical setting, in connection with increased model capacity.
研究の動機と目的
- 複雑な環境に複数の動く人間が存在する、大規模で多様で写真のようにリアルな3Dデータセットが、完全な3D真値を備えていないという課題に対処する。
- 豊富なアノテーションを備えたスケーラブルで高精細な合成データセットを構築することで、3D人間ポーズおよびシェイプ推定におけるシミュレーションから現実へのギャップを埋める。
- 多様なボディシェイプ、動き、シーンの条件にわたる一般化を可能にするモデルの開発と評価を可能にする。
- 弱い教師付き学習と合成データを用いたモデルの訓練を支援し、実世界のベンチマークでの性能向上を実現する。
- 多様で衝突のない、視覚的にリアルな人間のアニメーションを、複雑なシーンに自動で生成できるスケーラブルで自動化されたパイプラインを提供する。
提案手法
- 100体の衣装を着た多様な人物の3Dスキャンを、パrametricなシェイプ変動を伴うGHUM 3Dボディモデルに適合させる。
- CMU Mocapから取得した100本の実際の人間のモーショングラフデータを用いて、適合したGHUMモデルのリターゲティングとアニメーション化を行う。
- 100の合成された屋内・屋外シーンに、リアルなライティングとオクルージョンを備えた複数のアニメートされた人間を自動で配置する。
- 高精細なゲームエンジンを用いて、4K/HDRの画像および動画を、一貫したカメラモーションと多様な視点でレンダリングする。
- シーン配置中に人間同士、および人間と環境との間で衝突を自動回避する仕組みを実装する。
- 3Dポーズ/シェイプ、人間セグメンテーション、ボディパーツの局所化、時間的対応関係を含む、豊富なアノテーションを生成する。
実験結果
リサーチクエスチョン
- RQ1完全な3D真値を備えた大規模な合成データは、実世界のシナリオにおける3D人間ポーズおよびシェイプ推定を向上させ得るか?
- RQ2弱い教師付きの実データと組み合わせた場合、合成データはシミュレーションから現実へのドメインギャップをどの程度効果的に低減できるか?
- RQ3合成データと実データの混合学習において、モデル容量を増加させることで性能がどの程度向上するか?
- RQ4ボディシェイプ、動き、シーンの複雑さの多様性が、モデルの一般化性および耐性に与える影響はどの程度か?
- RQ5自動的でスケーラブルなパイプラインは、正確な3Dアノテーションを備えた、衝突のない高精細な人間アニメーションを、複雑な環境で生成可能か?
主な発見
- HSPACEで学習させることで、3D人間ポーズおよびシェイプ推定の性能が顕著に向上し、Human3.6MベンチマークにおけるMPJPE-PAが39.0に低下し、以前のSOTAを上回った。
- THUNDRモデルをHSPACEデータで微調整した結果、MPJPE-PAは39.8から39.0に低下し、プロトコルP1下でMPJPE-Tは143.9から132.5に低下した。
- HITI(実データ)とHSPACE(合成データ)を混合して学習させたモデルは性能向上を示し、最大のT-THUNDRモデルではMPJPE-PAが47に低下した。
- HSPACEテストセットにおいて、モデル容量をSMALLからBIGに増加させたことで、MPJPE-PAは50から47に低下し、より大きなアーキテクチャの利点が示された。
- 完全な3D監視(S#3D)を備えた合成データの導入により、特に実データと弱い教師付き学習を組み合わせた場合に顕著な性能向上が得られた。
- このデータセットは、HSPACE-TESTおよびHITI-TESTセットでも強力な性能を発揮し、標準ベンチマークを越えた一般化性能の評価に有用であることが示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。