정민화 교수
Minhwa Chung
서울대학교 · 컴퓨터과학
연구실 소개
정민화 교수의 연구실은 신경계 질환의 조기 징후로 나타나는 음성 이상을 자동으로 평가하고 진단하는 데 초점을 맞추고 있습니다. 주로 프로소디(음성의 리듬, 높낮이, 속도) 분석을 기반으로 한 언어 장애 자동 평가 기술과 외국어 학습자, 장애가 있는 환자의 특수한 음성 특성에 맞는 음성 인식 및 평가 시스템을 개발하고 있습니다. 특히 한국어와 영어 데이터를 기반으로 한 다국어 음성 분석 기술과 인공지능 기반의 자동 진단 알고리즘 개발이 핵심입니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15One of the first cues for many neurological disorders are impairments in speech. The traditional method of diagnosing speech disorders such as dysarthria involves a perceptual evaluation from a trained speech therapist. However, this approach is known to be difficult to use for assessing speech impairments due to the subjective nature of the task. As prosodic impairments are one of the earliest cues of dysarthria, the current study presents an automatic method of assessing dysarthria in a range
This study acoustically examines the quality of fricatives produced by ten dysarthric speakers with cerebral palsy. Previous similar studies tend to focus only on sibilants, but to obtain a better understanding of how dysarthria affects fricatives we selected a range of samples with different places of articulation and voicing. The Universal Access (UA) Speech database was used to select thirteen words beginning with one of the English fricatives (/f/, /v/, /s/, /z/, /∫/, /ð/). The following fou
In this paper we will introduce the work of creation of a speech database to develop speech technology for disabled persons, which has been done as part of a national program to help better life for Korean people. We will report about the creation of speech database of a total of 160 persons: prompting items, designs, etc. for the creation of a database which is needed to develop an embedded key-word spotting speech recognition system tailored for the persons disabled in articulation. The create
This paper proposes a method for automatic pronunciation assessment of Korean spoken by L2 learners by selecting the best feature set from a collection of the most well-known features in the literature. The L2 Korean Speech Corpus is used for assessment modeling, where the native languages of the L2 learners are English, Chinese, Japanese, Russian, and Mongolian. In our system, learners' speech is forced-aligned and recognized using a native Korean acoustic model. Based on these results, various
This paper presents a parallel natural language processing system implemented on a marker-passing parallel AI computer, the Semantic Network Array Processor (SNAP). Our system uses a memory-based parsing approach in which parsing is viewed as a memory search process. Linguistic information is stored as phrasal patterns in a semantic network knowledge base distributed over the memory of the parallel computer. Parsing is performed by recognizing and linking phrasal patterns that reflect a sentence
Detection of children with autism spectrum disorder (ASD) based on speech has relied on predefined feature sets due to their ease of use and the capabilities of speech analysis. However, clinical impressions may not be adequately captured due to the broad range and the large number of features included. This paper demonstrates that the knowledge-driven speech features (KDSFs) specifically tailored to the speech traits of ASD are more effective and efficient for detecting speech of ASD children f
Massively parallel computers offer not only improved speed but also a new perspective on computer vision, production systems, neural networks, and other AI applications. However, not much work has been done to apply parallel processing to natural-language processing, even though most sequential natural-language systems slow down as knowledge bases grow to realistic sizes and as linguistic features are added to handle special cases. To demonstrate the potential of parallel systems for natural-lan
This paper examines variations of Korean segments produced by Japanese learners of Korean. For corpus-based statistical analysis, we have used Korean read speech corpus produced by Japanese learners. Contrastive analysis of the target language and the source language is performed to provide information for interpreting the results of corpus analysis. Segmental variations are analyzed by aligning canonical phonetic transcriptions with auditory phonetic transcriptions of the corpus. The results sh
This study focuses on the issue of automatic severity classification of dysarthric speakers based on speech intelligibility. Speech intelligibility is a complex measure that is affected by the features of multiple speech dimensions. However, most previous studies are restricted to using features from a single speech dimension. To effectively capture the characteristics of the speech disorder, we extracted features of multiple speech dimensions: voice quality, prosody, and pronunciation. Voice qu
This paper describes a method of building Korean conversational speech data in the emergency medical domain and proposes an annotation method for the collected data in order to improve speech recognition performance. To suggest future research directions, baseline speech recognition experiments were conducted by using partial data that were collected and annotated. All voices were recorded at 16-bit resolution at 16 kHz sampling rate. A total of 166 conversations were collected, amounting to 8 h
This paper proposes an economical and effective phonetic transcription method for dealing with a large amount of nonnative English speech corpus. The method provides a consistent transcription agreement, although the corpus is transcribed by non-natives. To minimize the possibility of confusion in transcription process, forced aligned phone sequences and a set of possible mispronunciation candidate phones that Korean L2 learners are expected to make are given to the Korean transcribers for refer
Presents a parallel memory-based parser called PARALLEL, which is implemented on a marker-passing parallel AI computer called the Semantic Network Array Processor (SNAP). In the PARALLEL memory-based parser, the parallelism in natural language processing is utilized by a memory search model of parsing. Linguistic information is stored as phrasal patterns in a semantic network knowledge base that is distributed over the memory of the parallel computer. Parsing is performed by recognizing and link
대표 연구 분야
정민화 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.