Skip to main content
QUICK REVIEW

[Paper Review] What Makes Kevin Spacey Look Like Kevin Spacey

Supasorn Suwajanakorn, Ira Kemelmacher-Shlizerman|arXiv (Cornell University)|Jun 2, 2015
Face recognition and analysis26 references4 citations
TL;DR

This paper presents a system that reconstructs a controllable 3D facial model from unstructured photo collections to capture a person's 'persona'—their distinctive appearance and behavior. By combining 3D face reconstruction, tracking, alignment, and multi-texture modeling, the method enables puppeteering: driving the reconstructed model with another person's video while preserving the target's identity, shape, and texture, achieving highly realistic and consistent results on celebrities using only Internet imagery.

ABSTRACT

We reconstruct a controllable model of a person from a large photo collection that captures his or her {\em persona}, i.e., physical appearance and behavior. The ability to operate on unstructured photo collections enables modeling a huge number of people, including celebrities and other well photographed people without requiring them to be scanned. Moreover, we show the ability to drive or {\em puppeteer} the captured person B using any other video of a different person A. In this scenario, B acts out the role of person A, but retains his/her own personality and character. Our system is based on a novel combination of 3D face reconstruction, tracking, alignment, and multi-texture modeling, applied to the puppeteering problem. We demonstrate convincing results on a large variety of celebrities derived from Internet imagery and video.

Motivation & Objective

  • To develop a system that captures a person’s 'persona'—their unique visual identity and behavioral traits—using only unstructured photo and video collections.
  • To enable realistic puppeteering of one person (B) using video of another (A), such that B mimics A’s expressions while retaining B’s own facial identity and texture.
  • To overcome the limitations of prior methods that require controlled scans or lab sessions by using large-scale, unstructured Internet imagery for 3D model reconstruction.
  • To synthesize high-fidelity, expression-dependent textures from diverse photos, preserving wrinkles, creases, and lighting variations.
  • To demonstrate the feasibility of creating realistic, controllable avatars of well-photographed individuals without requiring active participation or specialized equipment.

Proposed method

  • Reconstructs a 3D face model of the target person (B) from a large photo collection using 3D face reconstruction and shape deformation techniques.
  • Aligns and tracks facial geometry and motion from the driver’s video (A) using 3D optical flow to estimate expression-driven deformations.
  • Transfers the motion from the driver (A) to the puppet (B) by deforming B’s reconstructed 3D shape using the estimated motion field.
  • Synthesizes detailed, expression-dependent textures for B by multi-scale blending of aligned photos from the collection, preserving high-frequency details like wrinkles.
  • Uses a similarity measure based on facial expression, lighting, and appearance to weight photos during texture averaging, ensuring consistency and realism.
  • Applies a hierarchical blending process across image pyramids to maintain sharpness and reduce flicker, especially under varying lighting conditions.

Experimental results

Research questions

  • RQ1What aspects of a person’s appearance and behavior define their unique 'persona' in a way that allows recognition across diverse roles?
  • RQ2Can a realistic, controllable 3D facial model be reconstructed from unstructured, uncalibrated photo collections without scanning or lab sessions?
  • RQ3How can facial motion from one person be transferred to another while preserving the target’s identity, shape, and texture?
  • RQ4What techniques enable consistent, high-fidelity texture synthesis from large, diverse photo collections with varying lighting and expressions?
  • RQ5To what extent can the system generalize to unseen expressions or lighting conditions using only existing photo data?

Key findings

  • The system successfully generates highly realistic, expression-dependent textures from large photo collections, preserving fine details such as wrinkles and creases.
  • Results show that using actor B’s shape and texture with actor A’s motion produces more convincing and believable puppeteering than alternatives.
  • The method outperforms baseline approaches—static warping, unaligned weighted averaging, and prewarped averaging—by producing sharper, more consistent, and expression-aware textures.
  • Texture synthesis remains robust under varying lighting conditions, as low-frequency shading effects are averaged out, while high-frequency details are preserved when sufficient similar-expression photos are available.
  • The system achieves consistent results even when references have strong illumination or are in black-and-white, due to the multi-scale blending and similarity-based weighting.
  • Performance degrades with smaller collections or limited expression variation, but can be mitigated by adjusting similarity thresholds and standard deviation in the blending process.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.