[논문 리뷰] FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing
FineMoGen는 공간적 및 시간적 의존성을 명시적으로 모델링하기 위해 새로운 Spatio-Temporal Mixture Attention (SAMI) 메커니즘을 활용하는 확산 기반 프레임워크로, 제로샷 및 완전히 감독된 설정 모두에서 최신 기술 수준의 성능을 달성한다. 또한 대규모 언어 모델(Large Language Models, LLMs)을 통한 상호작용형 운동 편집을 가능하게 하여 텍스트 기반 운동 합성에서의 제어성과 사용자 상호작용을 크게 향상시킨다.
Text-driven motion generation has achieved substantial progress with the emergence of diffusion models. However, existing methods still struggle to generate complex motion sequences that correspond to fine-grained descriptions, depicting detailed and accurate spatio-temporal actions. This lack of fine controllability limits the usage of motion generation to a larger audience. To tackle these challenges, we present FineMoGen, a diffusion-based motion generation and editing framework that can synthesize fine-grained motions, with spatial-temporal composition to the user instructions. Specifically, FineMoGen builds upon diffusion model with a novel transformer architecture dubbed Spatio-Temporal Mixture Attention (SAMI). SAMI optimizes the generation of the global attention template from two perspectives: 1) explicitly modeling the constraints of spatio-temporal composition; and 2) utilizing sparsely-activated mixture-of-experts to adaptively extract fine-grained features. To facilitate a large-scale study on this new fine-grained motion generation task, we contribute the HuMMan-MoGen dataset, which consists of 2,968 videos and 102,336 fine-grained spatio-temporal descriptions. Extensive experiments validate that FineMoGen exhibits superior motion generation quality over state-of-the-art methods. Notably, FineMoGen further enables zero-shot motion editing capabilities with the aid of modern large language models (LLM), which faithfully manipulates motion sequences with fine-grained instructions. Project Page: https://mingyuan-zhang.github.io/projects/FineMoGen.html
연구 동기 및 목표
- 텍스트 기반 운동 생성에서 신체 부위에 대한 세밀한 공간적·시간적 동작을 지정할 수 없는 문제를 해결하기 위해.
- 자연어 지시어를 통한 제로샷 운동 편집을 가능하게 하여 사용자 상호작용과 적응성을 향상시키기 위해.
- 세밀한 공간적·시간적 애너테이션을 포함한 대규모 벤치마크를 구축하기 위해.
- 주의 메커니즘 내에서 공간적 및 시간적 구성 요소를 명시적으로 모델링하여 운동 생성 품질을 향상시키기 위해.
제안 방법
- 주의 메커니즘에서 공간적 및 시간적 모델링을 분리함으로써 글로벌 주의 템플릿 학습을 향상시키는 새로운 Spatio-Temporal Mixture Attention (SAMI) 메커니즘을 제안한다.
- 세밀한 특징을 적응적으로 추출하기 위해 희소 활성화된 Mixture-of-Experts (MoE)를 도입하여 특징 품질과 모델 효율성을 향상시킨다.
- 학습 및 추론 과정에서 단일 텍스트 기반 및 세밀한 다중부위 기반 기술을 모두 지원하여 조건 주입의 유연성을 확보한다.
- 대규모 언어 모델(Large Language Models, LLMs)을 통합하여 자연어 편집을 해석하고 세밀한 운동 기술을 자동으로 수정한다.
- 7개의 신체 부위와 운동 단계에 걸쳐 총 2,968개의 영상과 102,336개의 세밀한 공간적·시간적 애너테이션을 포함한 HuMMan-MoGen 데이터셋을 구축한다.
- 표준 벤치마크(HumanML3D, KIT-ML, BABEL)와 새로운 HuMMan-MoGen 데이터셋을 기반으로 제로샷 및 완전히 감독된 설정 모두에서 학습 및 평가를 수행한다.
실험 결과
연구 질문
- RQ1확산 모델은 세밀한 다중부위 공간적·시간적 기술을 정확히 반영하는 인간 운동 시퀀스를 생성할 수 있는가?
- RQ2주의 메커니즘이 운동 생성에서 공간적 및 시간적 의존성을 더 잘 모델링하기 위해 어떻게 재구조화될 수 있는가?
- RQ3세밀한 기술에 대한 완전한 감독 없이도 모델이 제로샷 운동 생성에 얼마나 잘 일반화되는가?
- RQ4대규모 언어 모델이 효과적으로 통합되어 자연어 기반의 상호작용형 운동 편집을 가능하게 할 수 있는가?
- RQ5명시적인 공간적·시간적 구성 모델링이 운동 품질과 일관성에 어떻게 기여하는가?
주요 결과
- FineMoGen는 제로샷 평가에서 HuMMan-MoGen 테스트 세트에서 R-Precision 점수 0.51을 기록하여 기준 모델들을 압도한다.
- 모델은 FID를 1.09로 감소시켜 기준 FID 2.87에 비해 운동 품질을 크게 향상시켰다.
- 절단 실험을 통해 공간적 및 시간적 독립성 구성 요소가 모두 필수적임을 확인하였으며, 전체 SAMI + MoE 구성에서 최고의 성능을 기록했다.
- LLM 통합으로 인해 효과적인 제로샷 운동 편집이 가능해져 사용자가 자연어 명령어를 통해 운동 시퀀스를 수정할 수 있게 되었다.
- HuMMan-MoGen 데이터셋은 102,336개의 세밀한 애너테이션을 포함하여 대규모 세밀한 운동 생성 연구를 위한 종합적인 벤치마크를 제공한다.
- 모델은 높은 다양성(5.71)과 다중 모odal 성질(5.17)을 유지하여 다양한 운동 스타일 간의 강건성과 일반화 능력을 입증했다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.