[论文解读] Decoding visemes: improving machine lipreading (PhD thesis).
本博士论文提出一种说话人相关的音素到视觉音素映射方法,以提升机器唇读的准确性,表明最优视觉音素集合(每位说话人11–35个)可显著提高分类性能。通过使用说话人特定的视觉音素聚类进行分层训练,并将结果解码回音素,该方法在现有方法基础上实现了显著的准确率提升,尤其相较于李氏映射(Lee’s map),后者被证明是最优基线。
Machine lipreading (MLR) is speech recognition from visual cues and a niche research problem in speech processing & computer vision. Current challenges fall into two groups: the content of the video, such as rate of speech or; the parameters of the video recording e.g, video resolution. We show that HD video is not needed to successfully lipread with a computer. The term viseme is used in machine lipreading to represent a visual cue or gesture which corresponds to a subgroup of phonemes where the phonemes are visually indistinguishable. A phoneme is the smallest sound one can utter, because there are more phonemes per viseme, maps between units show a many-to-one relationship. Many maps have been presented, we compare these and our results show Lee's is best. We propose a new method of speaker-dependent phoneme-to-viseme maps and compare these to Lee's. Our results show the sensitivity of phoneme clustering and we use our new knowledge to augment a conventional MLR system. It has been observed in MLR, that classifiers need training on test subjects to achieve accuracy. Thus machine lipreading is highly speaker-dependent. Conversely speaker independence is robust classification of non-training speakers. We investigate the dependence of phoneme-to-viseme maps between speakers and show there is not a high variability of visemes, but there is high variability in trajectory between visemes of individual speakers with the same ground truth. This implies a dependency upon the number of visemes within each set for each individual. We show that prior phoneme-to-viseme maps rarely have enough visemes and the optimal size, which varies by speaker, ranges from 11-35. Finally we decode from visemes back to phonemes and into words. Our novel approach uses the optimum range visemes within hierarchical training of phoneme classifiers and demonstrates a significant increase in classification accuracy.
研究动机与目标
- 为解决机器唇读中的高说话人依赖性问题,开发说话人特定的音素到视觉音素映射方法。
- 研究尽管语音真实值一致,不同说话人之间视觉音素轨迹的可变性。
- 确定每位说话人实现唇读分类准确率最大化的最优视觉音素数量。
- 通过使用说话人优化的视觉音素集合进行分层训练,提升传统机器唇读系统的性能。
- 利用改进的视觉音素分类方法,将视觉音素解码回音素和词汇。
- 评估先前音素到视觉音素映射的鲁棒性,特别是李氏映射,并提出一种优化的说话人相关替代方案。
提出的方法
- 提出一种新颖的说话人相关音素到视觉音素映射技术,与通用映射不同,该技术考虑了视觉音素轨迹中的个体说话人差异。
- 使用从每位说话人独特视觉音素聚类模式中提取的优化视觉音素集合,对音素分类器进行分层训练。
- 对比多种现有音素到视觉音素映射,发现李氏映射是性能最佳的基线,适合作为比较基准。
- 利用视觉音素解码重建音素和词汇,将优化后的视觉音素表示集成到最终识别流程中。
- 通过评估视觉音素集合的大小和在不同说话人之间的分布,分析音素聚类的敏感性。
- 使用说话人特定的训练数据验证该方法,表明高清视频(HD video)并非实现高性能的必要条件。
实验结果
研究问题
- RQ1与通用映射相比,说话人相关的视觉音素映射如何影响机器唇读的准确性?
- RQ2在机器唇读中,实现分类性能峰值的每位说话人最优视觉音素数量是多少?
- RQ3即使发出相同的音素,个体说话人之间的视觉音素轨迹差异有多大?
- RQ4使用优化视觉音素集合的分层训练能否显著提升音素分类准确率?
- RQ5李氏音素到视觉音素映射的性能与所提出的说话人相关方法相比如何?
主要发现
- 李氏音素到视觉音素映射在现有映射中表现最佳,是性能对比的最优基线。
- 每位说话人最优视觉音素数量在11至35之间,具体取决于个体发音特征。
- 即使发出相同的音素,不同说话人之间的视觉音素轨迹存在高度可变性,表明存在强烈的说话人依赖性。
- 先前的音素到视觉音素映射通常缺乏足够的视觉音素粒度,导致分类性能不理想。
- 将所提出的说话人相关视觉音素映射方法集成到传统机器唇读系统中,可显著提升分类准确率。
- 使用优化视觉音素集合的分层训练方法,可实现从视觉音素到音素和词汇的有效解码,从而整体提升系统性能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。