Skip to main content
QUICK REVIEW

[Paper Review] Speech animation using electromagnetic articulography as motion capture data

Ingmar Steiner, Korin Richmond|arXiv (Cornell University)|Oct 30, 2013
Phonetics and Phonology Research19 references3 citations
TL;DR

This paper presents a lightweight, motion capture-based method for animating speech articulators using electromagnetic articulography (EMA) data as input. By treating EMA coil positions as skeletal joints in a BVH-compatible workflow, the authors generate realistic 3D tongue and jaw animations with 95% correlation to measured EMA trajectories, enabling real-time visualization and integration into multimedia and audiovisual synthesis applications using open-source tools.

ABSTRACT

Electromagnetic articulography (EMA) captures the position and orientation of a number of markers, attached to the articulators, during speech. As such, it performs the same function for speech that conventional motion capture does for full-body movements acquired with optical modalities, a long-time staple technique of the animation industry. In this paper, EMA data is processed from a motion-capture perspective and applied to the visualization of an existing multimodal corpus of articulatory data, creating a kinematic 3D model of the tongue and teeth by adapting a conventional motion capture based animation paradigm. This is accomplished using off-the-shelf, open-source software. Such an animated model can then be easily integrated into multimedia applications as a digital asset, allowing the analysis of speech production in an intuitive and accessible manner. The processing of the EMA data, its co-registration with 3D data from vocal tract magnetic resonance imaging (MRI) and dental scans, and the modeling workflow are presented in detail, and several issues discussed.

Motivation & Objective

  • To enable real-time, intuitive visualization of speech articulation using EMA data as motion capture input.
  • To overcome limitations of rule-based or biomechanically complex models by leveraging existing EMA data directly.
  • To develop a lightweight, open-source pipeline for integrating articulatory animation into multimedia and AV speech synthesis systems.
  • To avoid cross-speaker normalization issues by co-registering EMA with subject-specific MRI and dental scans.
  • To demonstrate feasibility of using off-the-shelf 3D software for articulatory animation without custom simulation frameworks.

Proposed method

  • Processed EMA data from the mngu0 corpus as motion capture input using the BVH file format for compatibility with standard animation tools.
  • Created a 3D tongue mesh from high-resolution MRI and dental scans of the same speaker to ensure anatomical accuracy and subject-specificity.
  • Used spline inverse kinematics (IK) to deform a NURBS-based tongue model, with EMA coil positions driving control points on the spline.
  • Rigged the tongue using an armature system where vertex groups on the mesh were weighted to follow the deformed spline, enabling smooth deformation.
  • Applied a simple hinge-based modifier to animate the jaw based on EMA coil data on the mandibular incisors.
  • Performed manual co-registration of EMA, MRI, and dental scan data using the palate contour as a reference, ensuring spatial alignment.

Experimental results

Research questions

  • RQ1Can EMA data be effectively repurposed as motion capture data to drive realistic 3D articulatory animation using standard animation pipelines?
  • RQ2How accurately can EMA-driven animation preserve the kinematics of actual articulatory movements, especially given sparse mesh topology?
  • RQ3To what extent does the use of subject-specific MRI and dental scans improve the anatomical fidelity of the resulting animation compared to generic models?
  • RQ4What are the key limitations of using EMA coil positions directly as animation drivers, particularly regarding coil placement and data noise?
  • RQ5Can open-source, off-the-shelf tools produce a functional, real-time articulatory animation system without requiring custom biomechanical modeling?

Key findings

  • The EMA-driven animation achieved a mean correlation of 0.95 between measured EMA coil trajectories and corresponding vertices on the animated tongue mesh.
  • Despite the sparse topology of the tongue mesh, the animation preserved the overall kinematic characteristics of natural speech movements.
  • The method successfully avoided cross-speaker normalization issues by using subject-specific MRI and dental scans for model creation.
  • The use of BVH-compatible data and open-source tools (e.g., Blender) enabled a lightweight, reproducible, and integrable animation pipeline.
  • The approach demonstrated feasibility for real-time visualization and integration into audiovisual speech synthesis systems, though posture differences between MRI (supine) and EMA (upright) may affect tongue shape fidelity.
  • Several challenges remain, including manual segmentation of MRI data, error-prone manual co-registration, lack of collision detection, and sensitivity to initial bind pose and coil placement.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.