[Paper Review] Decoding visemes: improving machine lipreading (PhD thesis).
This PhD thesis proposes a speaker-dependent phoneme-to-viseme mapping method to improve machine lipreading accuracy, demonstrating that optimal viseme sets (11–35 per speaker) significantly enhance classification performance. By using hierarchical training with speaker-specific viseme clusters and decoding back to phonemes, the approach achieves a notable increase in accuracy over existing methods, particularly Lee’s map, which is shown to be the most effective baseline.
Machine lipreading (MLR) is speech recognition from visual cues and a niche research problem in speech processing & computer vision. Current challenges fall into two groups: the content of the video, such as rate of speech or; the parameters of the video recording e.g, video resolution. We show that HD video is not needed to successfully lipread with a computer. The term viseme is used in machine lipreading to represent a visual cue or gesture which corresponds to a subgroup of phonemes where the phonemes are visually indistinguishable. A phoneme is the smallest sound one can utter, because there are more phonemes per viseme, maps between units show a many-to-one relationship. Many maps have been presented, we compare these and our results show Lee's is best. We propose a new method of speaker-dependent phoneme-to-viseme maps and compare these to Lee's. Our results show the sensitivity of phoneme clustering and we use our new knowledge to augment a conventional MLR system. It has been observed in MLR, that classifiers need training on test subjects to achieve accuracy. Thus machine lipreading is highly speaker-dependent. Conversely speaker independence is robust classification of non-training speakers. We investigate the dependence of phoneme-to-viseme maps between speakers and show there is not a high variability of visemes, but there is high variability in trajectory between visemes of individual speakers with the same ground truth. This implies a dependency upon the number of visemes within each set for each individual. We show that prior phoneme-to-viseme maps rarely have enough visemes and the optimal size, which varies by speaker, ranges from 11-35. Finally we decode from visemes back to phonemes and into words. Our novel approach uses the optimum range visemes within hierarchical training of phoneme classifiers and demonstrates a significant increase in classification accuracy.
Motivation & Objective
- To address the high speaker dependence in machine lipreading by developing speaker-specific phoneme-to-viseme mappings.
- To investigate the variability in viseme trajectories across speakers despite consistent phonetic ground truth.
- To determine the optimal number of visemes per speaker for maximizing lipreading classification accuracy.
- To enhance conventional machine lipreading systems through hierarchical training using speaker-optimized viseme sets.
- To decode visemes back to phonemes and words using improved viseme-based classification.
- To evaluate the robustness of prior phoneme-to-viseme maps, especially Lee’s, and propose a refined speaker-dependent alternative.
Proposed method
- Proposes a novel speaker-dependent phoneme-to-viseme mapping technique, differing from generic maps by accounting for individual speaker variability in viseme trajectories.
- Employs hierarchical training of phoneme classifiers using the optimized viseme sets derived from each speaker’s unique viseme cluster patterns.
- Compares multiple existing phoneme-to-viseme maps, with Lee’s map identified as the best-performing baseline for comparison.
- Uses viseme decoding to reconstruct phonemes and words, integrating the refined viseme representation into the final recognition pipeline.
- Analyzes the sensitivity of phoneme clustering by evaluating viseme set size and distribution across speakers.
- Validates the method using speaker-specific training data, demonstrating that HD video is not necessary for high performance.
Experimental results
Research questions
- RQ1How does speaker-specific viseme mapping affect machine lipreading accuracy compared to generic maps?
- RQ2What is the optimal number of visemes per speaker for achieving peak classification performance in machine lipreading?
- RQ3To what extent do individual speakers vary in viseme trajectory even when producing the same phoneme?
- RQ4Can hierarchical training with optimized viseme sets significantly improve phoneme classification accuracy?
- RQ5How does the performance of Lee’s phoneme-to-viseme map compare to the proposed speaker-dependent method?
Key findings
- Lee’s phoneme-to-viseme map is the most effective among existing maps and serves as the best baseline for comparison.
- The optimal number of visemes per speaker varies between 11 and 35, depending on individual articulation patterns.
- There is high variability in viseme trajectories between speakers even when producing the same phoneme, indicating strong speaker dependence.
- Prior phoneme-to-viseme maps often lack sufficient viseme granularity, leading to suboptimal classification performance.
- The proposed speaker-dependent viseme mapping method significantly increases classification accuracy when integrated into a conventional machine lipreading system.
- The hierarchical training approach using optimized viseme sets enables effective decoding from visemes back to phonemes and words, enhancing overall system performance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.