[논문 리뷰] Decoding visemes: improving machine lipreading (PhD thesis).
이 박사학위논문은 기계적 입술 읽기 정확도를 햖스키기 위해 발화자에 의존하는 음소-비세미 맵핑 방법을 제안하며, 최적의 비세미 집합(발화자당 11~35개)이 분류 성능을 크게 향상시킨다는 것을 입증한다. 발화자별 비세미 클러스터를 고려한 계층적 훈련을 통해 음소로 복원하는 디코딩을 수행함으로써, 이 방법은 기존 방법들, 특히 리의 맵핑 방식에 비해 뚜렷한 정확도 향상을 이룬다. 이는 가장 효과적인 기준 기반으로 입증되었다.
Machine lipreading (MLR) is speech recognition from visual cues and a niche research problem in speech processing & computer vision. Current challenges fall into two groups: the content of the video, such as rate of speech or; the parameters of the video recording e.g, video resolution. We show that HD video is not needed to successfully lipread with a computer. The term viseme is used in machine lipreading to represent a visual cue or gesture which corresponds to a subgroup of phonemes where the phonemes are visually indistinguishable. A phoneme is the smallest sound one can utter, because there are more phonemes per viseme, maps between units show a many-to-one relationship. Many maps have been presented, we compare these and our results show Lee's is best. We propose a new method of speaker-dependent phoneme-to-viseme maps and compare these to Lee's. Our results show the sensitivity of phoneme clustering and we use our new knowledge to augment a conventional MLR system. It has been observed in MLR, that classifiers need training on test subjects to achieve accuracy. Thus machine lipreading is highly speaker-dependent. Conversely speaker independence is robust classification of non-training speakers. We investigate the dependence of phoneme-to-viseme maps between speakers and show there is not a high variability of visemes, but there is high variability in trajectory between visemes of individual speakers with the same ground truth. This implies a dependency upon the number of visemes within each set for each individual. We show that prior phoneme-to-viseme maps rarely have enough visemes and the optimal size, which varies by speaker, ranges from 11-35. Finally we decode from visemes back to phonemes and into words. Our novel approach uses the optimum range visemes within hierarchical training of phoneme classifiers and demonstrates a significant increase in classification accuracy.
연구 동기 및 목표
- 기계적 입술 읽기에서 높은 발화자 의존성을 해결하기 위해 발화자별 음소-비세미 맵핑을 개발한다.
- 일관된 음소 기준값이 존재함에도 불구하고 발화자 간 비세미 궤적의 변동성을 조사한다.
- 입술 읽기 분류 정확도를 극대화하기 위해 발화자당 최적의 비세미 수를 규명한다.
- 발화자 최적화된 비세미 집합을 사용한 계층적 훈련을 통해 전통적인 기계적 입술 읽기 시스템을 향상시킨다.
- 개선된 비세미 기반 분류를 통해 비세미를 다시 음소와 단어로 디코딩한다.
- 이전의 음소-비세미 맵핑의 강건성, 특히 리의 맵핑을 평가하고 개선된 발화자에 의존하는 대안을 제안한다.
제안 방법
- 개별 발화자에서 비세미 궤적의 변동성을 반영하여 일반 맵핑과 다름없는 새로운 발화자에 의존하는 음소-비세미 맵핑 기법을 제안한다.
- 각 발화자 고유의 비세미 클러스터 패턴에서 유도된 최적화된 비세미 집합을 사용해 음소 분류기의 계층적 훈련을 수행한다.
- 여러 기존의 음소-비세미 맵핑을 비교하며, 리의 맵핑이 비교 기준으로서 가장 우수한 성능을 보임을 확인한다.
- 비세미 디코딩을 통해 음소와 단어를 재구성하며, 개선된 비세미 표현을 최종 인식 파이프라인에 통합한다.
- 비세미 집합 크기와 발화자 간 분포를 평가함으로써 음소 클러스터링의 민감도를 분석한다.
- 발화자별 훈련 데이터를 사용해 방법을 검증하며, 고성능 달성을 위해 HD 영상이 반드시 필요하지 않음을 입증한다.
실험 결과
연구 질문
- RQ1일반 맵핑에 비해 발화자별 비세미 맵핑은 기계적 입술 읽기 정확도에 어떤 영향을 미치는가?
- RQ2기계적 입술 읽기에서 최고의 분류 성능을 달성하기 위해 발화자당 최적의 비세미 수는 얼마인가?
- RQ3동일한 음소를 발화할 때에도 개별 발화자 간 비세미 궤적의 변동은 어느 정도인가?
- RQ4최적화된 비세미 집합을 사용한 계층적 훈련은 음소 분류 정확도를 뚜렷이 향상시킬 수 있는가?
- RQ5리의 음소-비세미 맵핑 성능은 제안된 발화자에 의존하는 방법과 비교해 어떻게 다른가?
주요 결과
- 기존 맵핑 중에서 리의 음소-비세미 맵핑이 가장 효과적이며, 비교 기준으로서 가장 적합한 기반으로 기능한다.
- 최적의 비세미 수는 발화자마다 다르게 11~35개 사이로 변동하며, 개인의 발음 패턴에 따라 달라진다.
- 동일한 음소를 발화할 때도 발화자 간 비세미 궤적에 높은 변동성이 존재하여 강한 발화자 의존성이 있음을 시사한다.
- 기존의 음소-비세미 맵핑은 종종 비세미의 해상도가 부족하여 최적의 분류 성능을 내지 못한다.
- 제안된 발화자에 의존하는 비세미 맵핑 방법은 기존의 기계적 입술 읽기 시스템에 통합될 경우 분류 정확도를 뚜렷이 향상시킨다.
- 최적화된 비세미 집합을 사용한 계층적 훈련 방식은 비세미에서 다시 음소와 단어로의 효과적인 디코딩을 가능하게 하여 전체 시스템 성능을 향상시킨다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.