Skip to main content

Mujeen Sung

Kyung Hee University · 情報科学

研究室紹介

Professor Mujeen Sung's research lab specializes in biomedical natural language processing, with a focus on advancing named entity recognition (NER) and normalization (NEN) in scientific literature. The lab develops deep learning-based tools such as BERN and BERN2 to improve the accuracy and efficiency of extracting and linking biomedical entities like diseases, drugs, and genes from vast biomedical texts. Their work emphasizes multi-task learning, retrieval-augmented generation, and handling complex challenges such as overlapping entities and incomplete synonym sets. The lab also explores the impact of collaboration dynamics on scientific innovation, particularly during global health crises like the COVID-19 pandemic.

biomedical NLPnamed entity recognitionnamed entity normalizationlarge language modelsknowledge graph construction

Research Overview

Papers
39
Total Citations
806
Papers (5y)
27
Primary Field
情報科学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
27total
2021
2022
2023
2024
2025
Citations per year (5y)
421total
20212022202320242025

Selected Papers

15
1
Article|133 citations·2020
Biomedical Entity Representations with Synonym Marginalization
Mujeen Sung, Hwisang Jeon, Jinhyuk Lee, Jaewoo Kang
OA

Biomedical named entities often play important roles in many biomedical text mining tools. However, due to the incompleteness of provided synonyms and numerous variations in their surface forms, normalization of biomedical entities is very challenging. In this paper, we focus on learning representations of biomedical entities solely based on the synonyms of entities. To learn from the incomplete synonyms, we use a model-based candidate selection and maximize the marginal likelihood of the synony

Molecular BiologyBiochemistry, Genetics and Molecular Biology
2
Article|132 citations·2019
A Neural Named Entity Recognition and Multi-Type Normalization Tool for Biomedical Text Mining
Donghyeon Kim, Jinhyuk Lee, Chan Ho So, Hwisang Jeon, Minbyul Jeong, Yong-Hwa Choi, Wonjin Yoon, Mujeen Sung, Jaewoo Kang
SJR Q1IEEE AccessOA

The amount of biomedical literature is vast and growing quickly, and accurate text mining techniques could help researchers to efficiently extract useful information from the literature. However, existing named entity recognition models used by text mining tools such as tmTool and ezTag are not effective enough, and cannot accurately discover new entities. Also, the traditional text mining tools do not consider overlapping entities, which are frequently observed in multi-type named entity recogn

Molecular BiologyBiochemistry, Genetics and Molecular Biology
3
Article|117 citations·2022
BERN2: an advanced neural biomedical named entity recognition and normalization tool
Mujeen Sung, Minbyul Jeong, Yonghwa Choi, Donghyeon Kim, Jinhyuk Lee, Jaewoo Kang
SJR Q1BioinformaticsOA

In biomedical natural language processing, named entity recognition (NER) and named entity normalization (NEN) are key tasks that enable the automatic extraction of biomedical entities (e.g. diseases and drugs) from the ever-growing biomedical literature. In this article, we present BERN2 (Advanced Biomedical Entity Recognition and Normalization), a tool that improves the previous neural network-based NER tool by employing a multi-task NER model and neural network-based NEN models to achieve muc

Molecular BiologyBiochemistry, Genetics and Molecular Biology
4
Article|107 citations·2024
Improving medical reasoning through retrieval and self-reflection with retrieval-augmented large language models
Minbyul Jeong, Jiwoong Sohn, Mujeen Sung, Jaewoo Kang
SJR Q1BioinformaticsOA

SUMMARY: Recent proprietary large language models (LLMs), such as GPT-4, have achieved a milestone in tackling diverse challenges in the biomedical domain, ranging from multiple-choice questions to long-form generations. To address challenges that still cannot be handled with the encoded knowledge of LLMs, various retrieval-augmented generation (RAG) methods have been developed by searching documents from the knowledge corpus and appending them unconditionally or selectively to the input of LLMs

Artificial IntelligenceComputer Science
5
Article|71 citations·2021
Learning Dense Representations of Phrases at Scale
Jinhyuk Lee, Mujeen Sung, Jaewoo Kang, Danqi Chen
OA

Jinhyuk Lee, Mujeen Sung, Jaewoo Kang, Danqi Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.

Artificial IntelligenceComputer Science
6
Article|51 citations·2021
Pandemics are catalysts of scientific novelty: Evidence from COVID‐19
Meijun Liu, Yi Bu, Chongyan Chen, Jian Xu, Daifeng Li, Yan Leng, Richard B. Freeman, Eric T. Meyer, Wonjin Yoon, Mujeen Sung, Minbyul Jeong, Jinhyuk Lee
SJR Q1Journal of the Association for Information Science and TechnologyOA

Abstract Scientific novelty drives the efforts to invent new vaccines and solutions during the pandemic. First‐time collaboration and international collaboration are two pivotal channels to expand teams' search activities for a broader scope of resources required to address the global challenge, which might facilitate the generation of novel ideas. Our analysis of 98,981 coronavirus papers suggests that scientific novelty measured by the BioBERT model that is pretrained on 29 million PubMed arti

Sociology and Political ScienceSocial Sciences
7
Article|43 citations·2020
Answering Questions on COVID-19 in Real-Time
Jinhyuk Lee, Sean S. Yi, Minbyul Jeong, Mujeen Sung, Wonjin Yoon, Yonghwa Choi, Miyoung Ko, Jaewoo Kang
OA

The recent outbreak of the novel coronavirus is wreaking havoc on the world and researchers are struggling to effectively combat it. One reason why the fight is difficult is due to the lack of information and knowledge. In this work, we outline our effort to contribute to shrinking this knowledge vacuum by creating covidAsk, a question answering (QA) system that combines biomedical text mining and QA techniques to provide answers to questions in real-time. Our system also leverages information r

Artificial IntelligenceComputer Science
8
Preprint|17 citations·2020
Transferability of Natural Language Inference to Biomedical Question Answering
Minbyul Jeong, Mujeen Sung, Gangwoo Kim, Donghyeon Kim, Wonjin Yoon, Jaehyo Yoo, Jaewoo Kang
arXiv (Cornell University)OA

Biomedical question answering (QA) is a challenging task due to the scarcity of data and the requirement of domain expertise. Pre-trained language models have been used to address these issues. Recently, learning relationships between sentence pairs has been proved to improve performance in general QA. In this paper, we focus on applying BioBERT to transfer the knowledge of natural language inference (NLI) to biomedical QA. We observe that BioBERT trained on the NLI dataset obtains better perfor

Artificial IntelligenceComputer Science
9
Article|15 citations·2023
Chemical identification and indexing in full-text articles: an overview of the NLM-Chem track at BioCreative VII
Robert Leaman, Rezarta Islamaj, Virginia Adams, Mohammed Alliheedi, João Rafael Almeida, Rui Antunes, Robert Bevan, Yung‐Chun Chang, Arslan Erdengasileng, Matthew Hodgskiss, Ryuki Ida, Hyunjae Kim
SJR Q1DatabaseOA

The BioCreative National Library of Medicine (NLM)-Chem track calls for a community effort to fine-tune automated recognition of chemical names in the biomedical literature. Chemicals are one of the most searched biomedical entities in PubMed, and-as highlighted during the coronavirus disease 2019 pandemic-their identification may significantly advance research in multiple biomedical subfields. While previous community challenges focused on identifying chemical names mentioned in titles and abst

Molecular BiologyBiochemistry, Genetics and Molecular Biology
10
Preprint|15 citations·2020
Can pandemics transform scientific novelty? Evidence from COVID-19
Meijun Liu, Yi Bu, Chongyan Chen, Jian Xu, Daifeng Li, Yan Leng, Richard B. Freeman, Eric T. Meyer, Wonjin Yoon, Mujeen Sung, Minbyul Jeong, Jinhyuk Lee

Scientific novelty is important during the pandemic due to its critical role in generating new vaccines. Parachuting collaboration and international collaboration are two crucial channels to expand teams' search activities for a broader scope of resources required to address the global challenge. Our analysis of 58,728 coronavirus papers suggests that scientific novelty measured by the BioBERT model that is pre-trained on 29 million PubMed articles, and parachuting collaboration dramatically inc

DemographySocial Sciences
11
Article|13 citations·2025
PubMed knowledge graph 2.0: Connecting papers, patents, and clinical trials in biomedical science
Jian Xu, Chao Yu, Jiawei Xu, Vetle I. Torvik, Jaewoo Kang, Mujeen Sung, Min Song, Yi Bu, Ying Ding, Yi Bu, Ying Ding
SJR Q1Scientific DataOA

Papers, patents, and clinical trials are essential scientific resources in biomedicine, crucial for knowledge sharing and dissemination. However, these documents are often stored in disparate databases with varying management standards and data formats, making it challenging to form systematic and fine-grained connections among them. To address this issue, we construct PKG 2.0, a comprehensive knowledge graph dataset encompassing over 36 million papers, 1.3 million patents, and 0.48 million clin

Molecular BiologyBiochemistry, Genetics and Molecular Biology
12
Article|11 citations·2025
Rationale-Guided Retrieval Augmented Generation for Medical Question Answering
Jiwoong Sohn, Yein Park, Chanwoong Yoon, Sihyeon Park, Hyeon Seok Hwang, Mujeen Sung, Hyunjae Kim, Jaewoo Kang
OA

Jiwoong Sohn, Yein Park, Chanwoong Yoon, Sihyeon Park, Hyeon Hwang, Mujeen Sung, Hyunjae Kim, Jaewoo Kang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.

Artificial IntelligenceComputer Science
13
Article|11 citations·2020
Transferability of Natural Language Inference to Biomedical Question Answering
Minbyul Jeong, Mujeen Sung, Gangwoo Kim, Donghyeon Kim, Wonjin Yoon, Jaehyo Yoo, Jaewoo Kang
CLEF (Working Notes)

Biomedical question answering (QA) is a challenging task due to the scarcity of data and the requirement of domain expertise. Pre-trained language models have been used to address these issues. Recently, learning relationships between sentence pairs has been proved to improve performance in general QA. In this paper, we focus on applying BioBERT to transfer the knowledge of natural language inference (NLI) to biomedical QA. We observe that BioBERT trained on the NLI dataset obtains better perfor

Artificial IntelligenceComputer Science
14
Preprint|9 citations·2020
Pandemics are catalysts of scientific novelty: Evidence from COVID-19
Meijun Liu, Yi Bu, Chongyan Chen, Jian Xu, Daifeng Li, Yan Leng, Richard B. Freeman, Eric T. Meyer, Wonjin Yoon, Mujeen Sung, Minbyul Jeong, Jinhyuk Lee
PubMedOA

Scientific novelty drives the efforts to invent new vaccines and solutions during the pandemic. First-time collaboration and international collaboration are two pivotal channels to expand teams' search activities for a broader scope of resources required to address the global challenge, which might facilitate the generation of novel ideas. Our analysis of 98,981 coronavirus papers suggests that scientific novelty measured by the BioBERT model that is pretrained on 29 million PubMed articles, and

Statistics, Probability and UncertaintyDecision Sciences
15
Book Chapter|8 citations·2022
Data-Centric and Model-Centric Approaches for Biomedical Question Answering
Wonjin Yoon, Jaehyo Yoo, Sumin Seo, Mujeen Sung, Minbyul Jeong, Gangwoo Kim, Jaewoo Kang
SJR Q2Lecture notes in computer science
Artificial IntelligenceComputer Science

Research Areas

Artificial IntelligenceMolecular BiologyFood ScienceSociology and Political ScienceDemographyStatistics, Probability and Uncertainty

Mujeen Sungの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。