Yonsei University · Computer Science
Professor Hyunwoo Kim's research lab specializes in computational linguistics and natural language processing, with a strong focus on empathy and prosocial behavior in dialogue systems. The lab investigates how language models can understand and respond to emotional and socially sensitive content by leveraging implicit causality, emotion cause detection, and commonsense reasoning. A key direction involves developing large-scale, human-annotated datasets—such as ProsocialDialog—that ground responses in social norms and real-world rules-of-thumb to promote safer, more ethical, and more empathetic AI interactions. The lab also explores second language acquisition, particularly how learners integrate syntactic constructions and form-meaning pairings, using both experimental psycholinguistics and NLP tools.
Figures are computed from collected data and may differ slightly.
Hyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Empathy is a complex cognitive ability based on the reasoning of others' affective states. In order to better understand others and express stronger empathy in dialogues, we argue that two issues must be tackled at the same time: (i) identifying which word is the cause for the other's emotion from his or her utterance and (ii) reflecting those specific words in the response generation. However, previous approaches for recognizing emotion cause words in text require sub-utterance level annotation
Most existing dialogue systems fail to respond properly to potentially unsafe user utterances by either ignoring or passively agreeing with them. To address this issue, we introduce ProsocialDialog, the first large-scale multi-turn dialogue dataset to teach conversational agents to respond to problematic content following social norms. Covering diverse unethical, problematic, biased, and toxic situations, ProsocialDialog contains responses that encourage prosocial behavior, grounded in commonsen
Abstract Implicit causality (IC) is a well-known phenomenon whereby certain verbs appear to create biases to remention either their subject or object in a causal dependent clause. This study investigated to what extent Korean learners of English made use of IC information for predictive processing at a discourse level, and whether L2 proficiency played a modulating role in this process. Results from a visual-world eye-tracking experiment showed early use of IC information in both L1 and L2 liste
Abstract One of the important components in second language (L2) development is to produce clause-level units of form–meaning pairings or argument structure constructions. Based on the usage-based constructionist approach that language development entails an ability to use more diverse, more complex, and less frequent constructions, this study tested whether constructional diversity and complexity predict L2 learners’ writing proficiency. Using a natural language processing tool called the Const
Abstract This study investigated the effects of construction types on Korean-L1 English-L2 learners’ verb–construction integration in online processing by presenting the ditransitive and prepositional dative constructions and manipulating the verb’s association strength within these constructions. Results of a self-paced reading experiment showed that the L2 group spent longer times in the verb–construction integration in the postverbal complement region when processing the ditransitive construc
Theory of mind (ToM) evaluations currently focus on testing models using passive narratives that inherently lack interactivity. We introduce FANToM, a new benchmark designed to stress-test ToM within information-asymmetric conversational contexts via question answering. Our benchmark draws upon important theoretical requisites from psychology and necessary empirical considerations when evaluating large language models (LLMs). In particular, we formulate multiple types of questions that demand th
This study investigates the influence of the semantic heaviness of verbs (i.e., heavy or light verbs) and language proficiency on second language (L2) learners’ use of constructional information in a sentence‐sorting task and a corpus analysis. Previous studies employing a sentence‐sorting task demonstrated that advanced L2 learners sorted English sentences according to argument structure constructions rather than lexical verbs. However, these studies collapsed both heavy (e.g., cut, throw ) and
This study investigates how the strength of referential biases associated with implicit vs explicit causality predicates in Korean affects Korean-speaking learners’ reference choices in English. Sentence-completion experiments with Korean (Experiment 1a) and English (1b) native speakers showed that Korean speakers referred to the subject more following predicates with explicit vs implicit causality marking, whereas English speakers showed no difference in referential bias for the English transla
The present study investigated the effects of construction‐based instruction on Korean English as a foreign language (EFL) learners’ production of the English argument structure constructions. Within the theoretical framework of construction grammar (Goldberg, 1995), the authors presented college students with English constructions in a hierarchical network and provided contextually meaningful visual scenes in connection with the language input. Results from translation and guided writing tasks,
Abstract The constructionist approach holds that an argument structure construction, a conventionalized form–meaning correspondence of a sentence, allows language users to efficiently access sentential information. This study investigated whether increased sensitivity to constructional information would enable second language learners to efficiently fuse information from a verb and a construction during real‐time sentence processing. Based on their performance in an English sentence sorting task
Abstract This study investigated the unresolved issue of potential sources of heritage language attrition. To test contributing effects of three learner variables – age of second language acquisition, length of residence, and language input – on heritage children's lexical retrieval accuracy and speed, we conducted a real-time word naming task with 68 children (age 11–14 years) living in South Korea who spoke either Chinese or Russian as a heritage language. Results of regression analyses showed
Open papers in the app to read, cite, and organize with AI.