Sanghoun Song
Korea University · Social Sciences
About the Lab
Professor Sanghoun Song's research lab specializes in computational linguistics, natural language processing (NLP), and sociolinguistic analysis with a strong focus on Korean language technologies. The lab investigates syntactic structures, ideological discourse in North Korean media, context-dependent hate speech detection, and the application of large language models in education and clinical linguistics. It also develops and evaluates NLP tools using large-scale corpora and deep learning techniques, particularly for low-resource and sociopolitically sensitive languages like Korean.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15본 논문은 북한의 주요 언론 매체인 노동신문의 언어 자료를 정제하여 이념적 어휘를 추출하고 이념적 어휘의 사회문화적 양상을 살피고자 하였다. 대규모 언어 자료를 처리하기 위해 딥러닝 중에서 단어임베딩의 방법론으로 사용하여 워드투벡 모형으로 이념적 어휘의 빈도와 유사어를 추출하였다. 2012-2013년도와 2016-2017년으로 구분하여 김정은 정권의 북한 사회의 변화 양상을 고찰하였다. 김정은 정권은 북한식 사회주의 강성대국을 위해 대외적 인식을 담론화하고 대내적 체제 유지 전략을 꾀하는 것으로 파악되었다. 이념적 어휘를 통한 북한 사회문화의 변화 양상을 연구자 주관을 배제하여 객관적인 자료를 분석함으로써 살폈다는 데에 의의가 있다. 대규모 북한 언어 자료의 계량적 연구 방법론을 제시하였다는 점에서도 앞으로 북한 관련 연구에 여러 측면에서 기여하리라고 기대한다.
Syntactic studies make use of the minimally pairwise sentences as an argumentation tool, because the pairs allow us to pay attention to the constraints of interest. Likewise, it is helpful to use a set of minimal pairs in deep learning-based experiments for assessing the syntactic ability of neural language models. In this context, this study verifies whether the deep learning Korean model has the ability to properly distinguish the well-formed expressions and the corresponding ill-formed expres
본 연구는 자연어 처리 분야에서 비교적 많이 연구되지 않은 맥락 의존적 혐오 표현에 주목하여, 이를 평가하기 위한 데이터 세트를 구축하고 생성형 언어 모델의 성능을 검증했다. 데이터는 한국의 대표적 온라인 커뮤니티 디시인사이드에서 수집되었으며, 맥락-댓글 총 2,005쌍으로 구성된 데이터 세트 KOCOH(KOrean COntext-Dependent Hate speech)가 구축되었다. 데이터의 유형은 실제 맥락과 맥락 의존적 혐오 표현, 만들어진 맥락과 맥락 의존적 비혐오 표현, 실제 맥락과 맥락 의존적 비혐오 표현의 세 가지로 구분되었다. 여섯 개의 최신 생성형 언어 모델(GPT-4o, GPT-4o-mini, Claude-3.5-sonnet, Claude-3.5-haiku, Bllossom, Ko-Gemma-2)에 대한 성능 평가 결과, F1 점수는 평균 60.73%로, 성능이 아직 만족스러운 수준에 도달하지 못했음을 확인했다. 특히, 언어 모델들은 고차원적 맥락이나 은어를 포함한
Internally Headed Relative Clauses have received much attention since the early days of generative grammar. This study primarily addresses whether the IHRCs in Korean are homogeneous or not. This study utilized the Sejong Spoken Corpus (i.e., transcription of naturally occurring speeches) to understand IHRCs with crucial reference to the discourse context (exploring 60,378 instances of relativization). The corpus exploration subclassifies the Korean IHRCs into three subtypes, using the replaceme
챗GPT 출시 이후 지난 1년간 교육 현장에서 챗GPT에 대한 무분별한 기대 와 우려가 공존하였지만, 실제로 교육이 크게 퇴보하거나 개선된 것은 아니다. 이러한 관찰을 토대로 본고는 챗GPT를 둘러싼 과장된 기대와 우려가 교육 현 장에서 AI를 둘러싼 잘못된 오해를 낳고 있음을 지적한다. 오히려 챗GPT를 비 롯한AI 기술을유연하게활용할경우교육방법론을향상시킬수있는가능성 이커질수있다. 나아가본고는교육현장에서챗GPT를포함한AI 기술을통 합하는 방법에 대한 실질적인 대안을 제시한다. AI 리터러시의 핵심 교과로 코 딩 교육의 중요성을 강조하며, AI를 비판적으로 이해하고 책임감 있게 사용할 수 있도록 AI 윤리 교육이 필수적으로 병행되어야 함을 지적한다. 본고가 AI 기 술의 교육적 적용에 대한 깊이 있는 논의를 촉진할 수 있기를 기대한다.
When studying the nature of human language, we frequently ask ourselvesthe following question: Do native speakers agree with our judgmentsof the sentences in question? Many of us have encountered quitea few sentences which linguists report to be grammatical but whichnon-linguists find ungrammatical. Linguists try their best in their languageanalyses to accommodate the native speakers’ intuitions in a systematicway, but these efforts are mostly confined to the so-called‘informal’ method. A natura
The purpose of this paper is to examine feasibility of replacing humans with deep learning in nativeness judgments and figure out in which way to develop the model in order to reach the level of humans by comparing nativeness judgments by deep learning and humans on English data. The controlled items, composed of 210 sentences, are categorized into two types: well-formedness test (i.e., no syntactic violation) and plausibility (i.e., no awkwardness) test items, most of which are excerpted from p
This study builds up a methodological pipeline to compare the lexical difference between the South Korean data and the North Korean data using the recent techniques of natural language processing. Assuming that the Chosun-ilbo and Rodong-shinmun are the counterpart of each other, we created the word embedding models to compare the distributional properties. We collected the data published in the newspapers from 2015 to 2017, preprocessed the texts, and then ran the Word2Vec library with the mani
배경 및 목적: 최근의 딥러닝 자연어 처리는 언어단위를 수치 벡터로 변환하여 공간상에서 연산을 도모하는 임베딩 기술을 활용한다. 본 연구는 이 기법을 언어병리학 유창성장애 데이터에 적용하여 비유창성의 위치와 분포 특성을 파악하고자 하였다. 방법: 110명의 중학생 이상 청소년 및 말더듬 성인의 읽기발화(800음절) 데이터를 음소 단위로 분절한 뒤 수치 벡터로 변환하여 거리 연산을 수행하였다. Word2Vec을 활용하여 코사인 유사도를 측정하여 각 비유창성 유형 별 유사성을 도출하고, 또한 전체 데이터를 t-SNE 그림으로 모델을 시각화하여 제시하였다. 또한, 음소 환경을 분석하고자 파라다이스-유창성검사-II의 읽기발화와 세종 코퍼스 데이터를 비교 분석하였다. 결과: 첫 번째, 총 8개의 ND 유형(‘UR’, ‘I’, ‘H’, ‘R1’)과 AD 유형(‘URa’, ‘Ia’, ‘Ha’, ‘R1a’)은 .9 이상의 유사도로 근접하여 출현하였다. 두 번째, ND 유형과 AD 유형 간의 분포적
In this paper, we test a working hypothesis that there is a semantic co-occurrence restriction between an adjective and the preposition in its complement in English. In the process, we propose a statistical method to evaluate the semantic constraint. The hypothesis of this study is that if a group of adjectives form a cluster sharing a common meaning, then they tend to co-occur with the same preposition. In order to test the validity of the hypothesis, we make use of two language resources, name
본고는 영어 및 한국어의 딥러닝 모델을 활용하여 언어 연구를 하는 방법론에 대해서 소개한다. 딥러닝 언어모델은 언어 표현의 연쇄가 가지는 확률적 자연스러움을 학습하므로, 그 자연스러움에 반하는 이상 분포에 대해서는 민감하게 반응한다. 이러한 이상치를 계산하는 심리언어학적 방식이 surprisal이다. 이 산술식을 이용한 언어 연구는 사실상 언어의 전 층위에 적용 가능하다. 형태론, 통사론, 의미론 등의 문장 단위 구성은 물론이며 담화 및 정보구조 등의 연구에도 사용할 수 있다. 나아가 언어 데이터에 함축되어 있는 인간의 세계 지식 및 상식 판단에 대해서도 준용할 수 있다. 본고는 surprisal 기반 실험을 실시할 때 주요한 고려 사항에 대해서도 개괄한다. 물론, 딥러닝 기반 방법이 자연언어에 대한 모든 것에 해법을 줄 수 있는 만능열쇠는 아니다. 그러나 인간 언어를 분석하기 위한 새로운 도구로서 실효성을 가진다는 점에서 앞으로 그 활용 여지가 크다. 관심있는 연구자의 편의를 위해
본 연구는 자연어처리 기반 인공지능이 발달장애인을 위한 보조공학으로 기능할 능성을 제시하며, 향후 연구 및 기술 개발의 방향을 모색한다. 인공지능 기술이 발전을 거듭하면서, 다양한 장애 보조공학으로 활용되고 있다. 여러 인공지능 기술 가운데 자연어처리는 특히 발달장애인의 학습과 의사소통을 지원하는 데 유용할 수 있다. 본 연구는 자연어처리 기반 보조공학이 발달장애인에게 제공할 수 있는 세 가지 주요 이점을 논의한다. 첫째, 텍스트 분석의 방법을 통하여 저렴하면서도 신속한 진단/평가 도구로 활용될 수 있다. 둘째, 발달장애인의 일상적인 생활과 학습 과정에서 지속적인 지원 도구로 기능한다. 셋째, 언어적 패턴 분석을 통해 연구자의 분석 정확도를 향상하고, 데이터 기반의 객관적인 연구 환경을 조성할 수 있다. 넷째, 특수교육 현장에서 교사의 보조자로 기능하며, 발달장애 학생들의 사회적 상호작용을 촉진할 수 있다. 종합하자면, 자연어처리 환경은 멀티모달 인공지능에 대비하여 발달장애인의 조기
This study exploits the benefits of combining corpora and experimental methods to investigate how Korean speakers acquire unaccusativity in English. Focusing on overpassivization, we comprehensively investigated three questions: (i) Are Korean speakers sensitive to the unaccusative/unergative distinction in English? (ii) Are they able to distinguish unaccusatives from transitives? (iii) Which factors among agentivity, telicity, and animacy do they rely on? We explored both native and learner cor
This paper provides a corpus study on subjunctives in Korean in a way of comparative semantics. The whole arguments of this paper are bolstered by distributional evidence taken from naturally occurring bitexts (i.e. a bilingual corpus), in which one sentence in a language is aligned with one translation in the other language. Since previous studies regard past tense morphology as the main component to express irrealis and uncertainty, this paper accordingly checks out whether the past tense morp
The present study computes the selectional preference the verbal items show in relation to their co-occurring subjects and objects in Korean. The selectional preference indicates how significant is the semantic corelation between a verbal item and the class of nouns that appear in its argument position. The selectional preference measurements presented in the current work is automatically measured in a bottom-up way, using two types of language resources: One is a parsed corpus (the Sejong Korea
Research Areas
Dive deeper into Sanghoun Song's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.