[論文レビュー] Generalizable Neural Performer: Learning Robust Radiance Fields for Human Novel View Synthesis
本論文は Generalizable Neural Performer (GNR) を紹介します。 sparse views からの自由視点人間合成のための一般化可能な暗黙的放射場フレームワークで、暗黙幾何学的身体埋め込みとスクリーン空間の Occlusion-aware な外観ブレンディングを組み合わせ、ケースごとのファインチューニングなしに対象間および姿勢間のロバストなレンダリングを実現します。 また、ロバストな評価のための GeneBody-1.0 データセットを提示します。
This work targets at using a general deep learning framework to synthesize free-viewpoint images of arbitrary human performers, only requiring a sparse number of camera views as inputs and skirting per-case fine-tuning. The large variation of geometry and appearance, caused by articulated body poses, shapes and clothing types, are the key bottlenecks of this task. To overcome these challenges, we present a simple yet powerful framework, named Generalizable Neural Performer (GNR), that learns a generalizable and robust neural body representation over various geometry and appearance. Specifically, we compress the light fields for novel view human rendering as conditional implicit neural radiance fields from both geometry and appearance aspects. We first introduce an Implicit Geometric Body Embedding strategy to enhance the robustness based on both parametric 3D human body model and multi-view images hints. We further propose a Screen-Space Occlusion-Aware Appearance Blending technique to preserve the high-quality appearance, through interpolating source view appearance to the radiance fields with a relax but approximate geometric guidance. To evaluate our method, we present our ongoing effort of constructing a dataset with remarkable complexity and diversity. The dataset GeneBody-1.0, includes over 360M frames of 370 subjects under multi-view cameras capturing, performing a large variety of pose actions, along with diverse body shapes, clothing, accessories and hairdos. Experiments on GeneBody-1.0 and ZJU-Mocap show better robustness of our methods than recent state-of-the-art generalizable methods among all cross-dataset, unseen subjects and unseen poses settings. We also demonstrate the competitiveness of our model compared with cutting-edge case-specific ones. Dataset, code and model will be made publicly available.
研究の動機と目的
- Sparse multi-view inputs から per-case fine-tuning なしで任意の人間パフォーマーの自由視点画像を合成するという課題に対処する。
- ポーズ、形状、衣装の変化を扱える堅牢で一般化可能なニューラル身体表現を開発する。
- 幾何学的 priors とソースビューの外観手掛かりを組み込み、ビュー間での幾何学的忠実度と外観のリアリズムを向上させる。
- GeneBody-1.0 による多様なマルチビュー・データセットを提供し、一般化可能な人間レンダリングを評価する。
提案手法
- Implicit Geometric Body Embedding を導入し、SMPLx ベースの幾何とマルチビュー手掛かりでニューラル放射場を条件づける。
- SMPLx 表面に基づく有符号距離関数(SDF)と canonical-space セマンティック埋め込みを用いて身体幾何をアンカーする。
- マルチビュー画像特徴を身体埋め込みと抽出・融合し、条件付き NeRF フレームワークで放射場を情報づける。
- Screen-Space Occlusion-Aware Appearance Blending (SSOA-AB) を提案し、ブレンディングをオクルージョンマップとビューアテンションベースのブレンディングに分解することで、ソースビューのテクスチャが放射予測を修正しつつ、未観測領域でのゴースティングを抑制する。
- 体積レンダリングを用いた条件付き NeRF でレンダリングし、3D 真値が利用可能な場合には占有・オクルージョン監督を含む写真測量・幾何学的損失で訓練する。
- GeneBody-1.0 および ZJU-Mocap で訓練・評価し、一般化ベースライン(pixelNerf、IBRNet)およびケース固有法(NeuralBody、NeuralTexture、NHR、NeuralVolumes)と比較する。
実験結果
リサーチクエスチョン
- RQ1単一モデルが一般化可能で堅牢な暗黙の身体表現を学習し、個別対象のファインチューニングなしに任意の人間パフォーマーの新規ビューを高品質に合成できるか。
- RQ2幾何 priors とマルチビューのソースヒントを統合して、ポーズや衣服の変化に対する頑健性を向上させるにはどうすればよいか。
- RQ3スクリーン空間オクルージョン対応ブレンディングは、ゴースティングのない外観と sparse 入力下でのマルチビュー整合性を実現できるか。
- RQ4提案手法は unseen subjects および unseen poses に対して、state-of-the-art な一般化法やケース固有法と比較してどのように性能を発揮するか。
主な発見
- GNR は unseen IDs および unseen poses に対して一般化ベースラインと比較して GeneBody-1.0 および ZJU-Mocap で優れたレンダリング品質を達成する。
- GNR は困難な衣服やポーズ下でも堅牢な幾何整合と高品質な外観を示し、いくつかの最先端手法を上回る。
- アブレーション研究により、Implicit body embedding、Attention-based appearance blending、Screen-space occlusion-aware blending がそれぞれ幾何、レンダリング、外観忠実度の改善に寄与することが示される。
- 合成 RenderPeople データで、GNR はベースラインと比較して3D幾何再構成(Chamfer 距離の低下)と高い画像品質(PSNR/SSIM)を示す。
- GeneBody-1.0 データセットは、多様な被験体、ポーズ、衣服を提供し、現実世界に近いシナリオでの一般化可能な人間レンダリングを評価する。
- body embedding がない場合、attention がない場合、または occlusion-aware blending がない場合、GNR の性能は低下し、それぞれの成分の価値を強調する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。