Sangah Lee
Seoul National University · Computer Science
About the Lab
Professor Sangah Lee's research lab specializes in natural language processing (NLP) for low-resource and morphologically complex languages, with a strong focus on Korean. The lab develops efficient, linguistically informed NLP models—such as KR-BERT—by leveraging subword tokenization, optimized vocabulary design, and domain-specific data curation. It also explores the integration of language, literature, and music to foster holistic cognitive and emotional development through interdisciplinary education research.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15To evaluate the interactive effect of methylenetetrahydrofolate reductase (MTHFR) genotype and dietary factors on the development of breast cancer, a hospital based case-control study was conducted in South Korean study population consisting of 189 histologically confirmed incident breast cancer cases and their 189 age-matched controls without present or previous history of cancer. A PCR-RFLP method was used for the genotyping of MTHFR (C677T) and statistical evaluations were performed by uncond
최근 자연어처리에서 문장 단위의 임베딩을 위한 모델들은 거대한 말뭉치와 파라미터를 이용하기 때문에 큰 하드웨어와 데이터를 요구하고 학습하는 데 시간이 오래 걸린다는 단점을 갖는다. 따라서 규모가 크지 않더라도 학습 데이터를 경제적으로 활용하면서 필적할만한 성능을 가지는 모델의 필요성이 제기된다. 본 연구는 음절 단위의 한국어 사전, 자소 단위의 한국어 사전을 구축하고 자소 단위의 학습과 양방향 WordPiece 토크나이저를 새롭게 소개하였다. 그 결과 기존 모델의 1/10 사이즈의 학습 데이터를 이용하고 적절한 크기의 사전을 사용해 더 적은 파라미터로 계산량은 줄고 성능은 비슷한 KR-BERT 모델을 구현할 수 있었다. 이로써 한국어와 같이 고유의 문자 체계를 가지고 형태론적으로 복잡하며 자원이 적은 언어에 대해 모델을 구축할 때는 해당 언어에 특화된 언어학적 현상을 반영해야 한다는 것을 확인하였다.
Since the appearance of BERT, recent works including XLNet and RoBERTa utilize sentence embedding models pre-trained by large corpora and a large number of parameters. Because such models have large hardware and a huge amount of data, they take a long time to pre-train. Therefore it is important to attempt to make smaller models that perform comparatively. In this paper, we trained a Korean-specific model KR-BERT, utilizing a smaller vocabulary and dataset. Since Korean is one of the morphologic
User-generated texts include various types of stylistic properties, or noises. Such texts are not properly processed by existing morpheme analyzers or language models based on formal texts such as encyclopedias or news articles. In this paper, we propose a simple morphologically tight-fitting tokenizer (K-MT) that can better process proper nouns, coinages, and internet slang among other types of noise in Korean user-generated texts. We tested our tokenizer by performing classification tasks on K
This study aims to reveal convergent characteristics and creativity of Eco-poetry. Existing studies on convergence of science and art understand nature as a subject which has an objective truth understandable through scientific methodology. It has been identified that this point of view is opposed to the world-view of art which focused on aesthetic judgement and figuration of nature. There have been some attempts recently to blur such conflictual distinction and communicate across the borderline
본고는 학습자의 사고와 정서 발달에서 필수 불가결한 문화 양식인 문학과 음악에 대하여 범교과적 접근의 가능성을 살피고 연계 방향을 논의하였다. 이들은 매재 및 구성과 소통 방식, 특히 추상적 상징 체계인 언어의 개입 여부에 따라 차이를 지니게 되었지만, 형상적 사유와 상상력을 바탕으로 정서를 자극한다는 점에서 미적인 속성을 공유한다. 문학과 음악을 통한 미적 체험은 인간으로 하여금 잃어버렸던 총체성을 회복하게 한다는 점에서 주관적인 만족에 그치지 않고 사회를 긍정적으로 변화시킬 수 있는 주체를 형성하는 데 기여한다. 이러한 전인적 인간으로의 성장은 교육의 궁극적인 목표이기도 하다. 본고는 위와 같은 목표의 실현에 문학과 음악의 범교과적 접근이 기여할 수 있다는 전제 아래 교육과정 및 교과서의 분석을 통해 그 방향을 가늠해 보았다. 문학과 음악은 본질적으로 시간적 경험의 특성을 지니기 때문에 감상의 진행 과정에서 상호작용하며 미적 체험을 형성할 수 있다. 이는 문학을 음성으로 구현하는
대규모 코퍼스에 기반한 사전학습모델인 BERT 모델은 언어 모델링을 통해 텍스트 내의 다양한 언어 정보를 학습할 수 있다고 알려져 있다. 여기에는 별도의 언어 자질이 요구되지 않으나, 몇몇 연구에서 특정한 언어 지식을 추가 반영한 BERT 기반 모델이 해당 지식과 관련된 자연어처리 문제에서 더 높은 성능을 보고하였다. 본 연구에서는 감정 분석 성능을 높이기 위한 방법으로 한국어 감정 사전에 주석된 감정 극성과 강도 값을 이용해 감정 자질 임베딩을 구성하고 이를 보편적 목적의 BERT 모델과 결합하는 외적 결합과 지식 증류 방식을 제안한다. 감정 자질 모델은 작은 스케일의 BERT 모델을 적은 스텝 수로 학습하여 소요 시간과 비용을 줄이고자 했으며, 외적 결합된 모델들은 영화평 분류와 악플 탐지문제에서 사전학습모델의 단독 성능보다 향상된 결과를 보였다. 또한 본 연구는 기존의 BERT 모델 구조에 추가된 감정 자질이 언어 모델링 및 감정 분석의 성능을 개선시킨다는 것을 관찰하였다.
본고의 목표는 현대시의 운율을 시 교육에서 활용하기 위한 읽기 방법을 모색하는 데 있다. 운율은 개념 자체가 지닌 혼란에 더불어 국내로 수용되면서 서구의 언어 및 시사(詩史)와의 차이로 인해 명확한 합의가 어려운 용어가 되었다. 본고는 운율 교육의 목표가 지식을 ‘아는 것’이라기보다 이를 활용하여 시를 읽고 해석할 수 있는 능력 함양에 있다는 데 주목하였다. 따라서 운율을 느끼기 위해 학습자의 시 읽기는 음독 상황을 뜻하는 것이어야 하며, 이러한 음독의 과정에서 시의 전체적 의미와 운율 간의 관계를 찾을 수 있어야 한다는 조건 아래 연구를 진행하였다. 길이와 형태 측면에서 자유로운 확대를 꾀한 현대시 텍스트를 음독하며 학습자들은 의도된 호흡의 패턴을 느낌으로써 시의 의미를 추론해 나갈 수 있다. 이러한 경험은 이후 시의 다양한 운율에 대한 주체적 발견과 해석을 하는 데 유용한 도움을 줄 수 있을 것이다.
In the present study, three-dimensional aerodynamic optimization of high pressure turbine nozzle for turbofan engine was performed. For this, Kriging surrogate model was built and refined iteratively by supplying additional experimental points until the surrogate model and CFX result has effective difference on objective function. When the surrogate model satisfied this reliability condition and developed enough, optimum point was investigated. Commercial program PIAnO was used for optimization
Copper indium gallium sulfur selenide (Cu(In<sub>1-x</sub>Ga<sub>x</sub>)SeS, CIGS) thin film solar cells are fabricated using a solution-based process, and their defect models are studied through a computer-aided design method. Cu(In<sub>1-x</sub>Ga<sub>x</sub>)SeS is structured with a graded bandgap by controlling the ambient gas and precursor composition, during the fabrication process. The defects in the CIGS are modeled as two donor-like defects, which are differently distributed as per the
The purpose of this study was to develop a computer software program for nutritional assessment using a Semi Quantitative Food Frequency Questionnaire (SQFFQs) and the 24-hour Recall Method. The software for the SQFFQ was divided into input, output, and database. For dietary analyses, recipe and food databases were used. The recipe database included 25 items and the food database was divided into 18 food groups. The food database was composed of 19 general nutrient items, 33 fatty acids, and 18
본고는 학습자들이 경제 교육을 통해 습득하는 이론과 지식은 실제 경제 현상 속에서 마주치게 되는 문제를 창의적으로 해결하는 데 기여해야 한다는 것에 중점을 두고 있다. 이를 위해 이론과 실제를 효과적으로 이어줄 수 있는 요소로서 ‘세대’ 개념에 초점을 맞추어 경제주체가 경제 현상 안에서 영향을 주고받는 양상을 살펴보고 이를 경제 교육에 적용할 수 있는 방법을 탐색하였다. 세대 개념 및 명칭은 그 구분의 비고정성으로 인한 분류의 다양성, 영역 복합적 성격에 기인한 입체성이라는 특성을 지닌다. 따라서 경제 기사 속 등장하는 경제 주체에게 하나의 세대 명칭이 부여되고 이것이 사용되는 양상은 세대의 분류나 명칭 선택에 개입하는 주체의 의도에 대해 인식하게 한다. 이를 추론함으로써 경제 교육의 장면에서는 비판적 사고 활동이 가능해지고 나아가 그 의도에 따른 사용역의 변화 역시 교수-학습의 내용으로 가능하게 된다. 이와 같은 활동을 통해 학습자들은 경제 주체가 경제 현상에서 어떠한 위치를 점하
In this study, an aero-acoustic analysis around pantograph of a high speed train is performed. Computational technique and grid system is validated with wind tunnel test result and unsteady acoustic pressure data are used for analyzing noise level of each part of pantograph. FLUENT is used for flow analysis and LES(Large Eddy Simulation) is applied for analyzing turbulent flow. For acoustic analysis, Ffowcs Williams-Hawkings(FW-H) acoustics model is used and it bring the aero-acoustic characteri
최근에는 계약서를 포함한 법률 문서들을 대량으로, 빠르고 정확하게 처리하기 위하여 인공지능을 활용한 자동화된 분석 방법이 요구된다. 계약서는 그 안에 필수적인 조항들이 모두 포함되었는지, 어느 한 쪽에 불리한 조항은 없는지 등을 확인하여 적격성을 검증할 수 있다. 이때 계약서를 이루는 조항들은 계약서의 종류와 관계없이 매우 정형적이고 반복적인 경우가 많다. 본 연구에서는 이러한 성격을 이용하여 계약서 내 조항별 분류 모델을 구축하였으며, 계약서의 관습적인 요구사항에 기반하여 구성한 키워드 임베딩을 구축하고 이를 BERT 임베딩과 결합하여 사용한다. 이때 BERT 모델은 한국어 사전학습모델을 법률 도메인 문서를 이용하여 미세 조정한 것이다. 각 조항의 분류 결과는 정확도 90.57과 90.64, F1 점수 93.27과 93.26으로 우수한 수준이며, 이렇게 계약서를 이루는 각 조항이 어떤 필수조항에 해당되는지의 예측 결과를 통해 계약서의 적격성을 검증할 수 있다.
Research Areas
Dive deeper into Sangah Lee's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.