Nan Hoesung
Korea University · Psychology
About the Lab
Professor Nan Hoesung's research lab specializes in speech processing, computational phonetics, and second language acquisition, with a strong focus on modeling speech production through articulatory gestures and task dynamics. The lab develops advanced computational systems—such as the TADA platform—to simulate and analyze the temporal coordination of speech gestures, enabling synthesis and annotation of natural speech at a gestural level. A key research direction involves creating data-driven, analysis-by-synthesis frameworks for automatic gestural annotation of speech, enhancing both speech recognition and second language learning applications. The lab also investigates how different types of input and instruction affect implicit linguistic knowledge, particularly in the context of idiom acquisition and processing.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15A portable computational system called TADA was developed for the Task Dynamic model of speech motor control [Saltzman and Munhall, Ecol. Psychol. 1, 333–382 (1989)]. The model maps from a set of linguistic gestures, specified as activation functions with corresponding constriction goal parameters, to time functions for a set of model articulators. The original Task Dynamic code was ported to the (relatively) platform-independent MATLAB environment and includes a MATLAB version of the Haskins ar
Timed grammaticality judgment tests (TGJT) and oral elicited imitation tests (OEIT) are considered reliable and valid measures of implicit linguistic knowledge, but studies consistently observe better performances on the TGJT than the OEIT due to the different types of processing they require: comprehension for the TGJT and production for the OEIT. This study examines whether degree of access to implicit knowledge is a function of processing type. Results from a series of factor analyses suggest
Speech can be represented as a constellation of constricting vocal tract actions called gestures, whose temporal patterning with respect to one another is expressed in a gestural score. Current speech datasets do not come with gestural annotation and no formal gestural annotation procedure exists at present. This paper describes an iterative analysis-by-synthesis landmark-based time-warping architecture to perform gestural annotation of natural speech. For a given utterance, the Haskins Laborato
Abstract This study explores the relevance and effectiveness of processing instruction in second language (L2) idiom learning by examining (1) whether structured input (SI) is more effective than non-SI, comprehension-based activities and (2) whether explicit information (EI) in addition to SI can facilitate L2 idiom learning. One hundred adult L2 English speakers were randomly assigned to one of six conditions: four groups who participated in SI activities in one of four EI conditions (i. e., n
This study investigates the fine-tuning of large-scale Automatic Speech Recognition (ASR) models, specifically OpenAI’s Whisper model, for domain-specific applications using the KsponSpeech dataset. The primary research questions address the effectiveness of targeted lexical item emphasis during fine-tuning, its impact on domain-specific performance, and whether the fine-tuned model can maintain generalization capabilities across different languages and environments. Experiments were conducted u
Speech can be represented as a constellation of constricting events, gestures, which are defined at distinct vocal tract sites, in the form of a gestural score. Gestures and their output trajectories, tract variables, which are available only in synthetic speech, have recently been shown to improve automatic speech recognition (ASR) performance. In this paper we propose an iterative analysis-by-synthesis landmark based time-warping architecture to obtain gestural scores for natural speech. Given
Previous work has shown that velar stops are produced with a forward movement during closure, forming a forward (anterior) loop for a VCV sequence, when the preceding vowels are back or mid. Are listeners aware of this aspect of articulatory dynamics? The current study used articulatory synthesis to examine how such kinematic patterns are reflected in the acoustics, and whether those acoustic patterns elicit different goodness ratings. In Experiment I, the size and direction of loops was modulat
This study investigates how second language (L2) listeners’ perception is affected by two factors: the listeners’ experience with the target dialect – North American English (NAE) vs. Standard Southern British English (SSBE) – and talkers’ language background: native vs. non-native talkers; i.e. interlanguage speech intelligibility benefit (ISIB) talker effects. Two groups of native-Korean-speaking listeners with different target English dialects – L1-Korean listeners of English as a second lang
End-to-end (E2E) automatic speech recognition (ASR) has achieved promising performance gains with the introduced self-attention network, Transformer. However, due to training time and the number of hyperparameters, finding the optimal hyperparameter set is computationally expensive. This paper investigates the impact of hyperparameters in the Transformer network to answer two questions: which hyperparameter plays a critical role in the task performance and training speed. The Transformer network
This study presents a model of automatic speech recognition (ASR) that is designed to diagnose pronunciation issues in children with speech sound disorders (SSDs) to replace manual transcriptions in clinical procedures. Because ASR models trained for general purposes mainly predict input speech into standard spelling words, well-known high-performance ASR models are not suitable for evaluating pronunciation in children with SSDs. We fine-tuned the wav2vec2.0 XLS-R model to recognise words as the
Speech can be represented as a set of discrete vocal tract constriction gestures (gestural score) defined at functionally distinct speech organs [tract variables (TVs)]. Using such gestures as sub-word units in an ASR system, variation in speech arising from coarticulation and reduction can be addressed. Since there is a lack of test corpora annotated with gestural scores, we develop a semi-automatic procedure for estimating and annotating gestural scores from natural speech databases using the
Research Areas
Dive deeper into Nan Hoesung's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.