Skip to main content

Sang-Young Kim

Ewha Womans University · 情報科学

研究室紹介

Professor Sang-Young Kim's research lab specializes in advancing artificial intelligence with a focus on multimodal learning, conversational AI, and the safety and robustness of large language models. The lab develops innovative deep learning and machine learning frameworks that enhance diagnostic accuracy in healthcare through multimodal data fusion, such as combining medical imaging and audiometric data. It also pioneers efficient, low-resource methods for semantic retrieval in dialogue systems, exemplified by the HEISIR framework, while critically investigating vulnerabilities in long-context LLMs to improve model safety. The lab's work bridges practical applications in healthcare and natural language processing with foundational research in AI reliability and interpretability.

multimodal learningconversational AIlarge language modelssemantic retrievalAI safety

Research Overview

Papers
25
Total Citations
26
Papers (5y)
25
Primary Field
情報科学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
25total
2022
2023
2024
2025
2026
Citations per year (5y)
26total
20222023202420252026

Selected Papers

15
1
Article|11 citations·2023
Toward Better Ear Disease Diagnosis: A Multi-Modal Multi-Fusion Model Using Endoscopic Images of the Tympanic Membrane and Pure-Tone Audiometry
T.W. Kim, Sangyeop Kim, Jaeyoung Kim, Yeonjoon Lee, June Choi
SJR Q1IEEE AccessOA

Chronic otitis media is characterized by recurrent infections, leading to serious complications, such as meningitis, facial palsy, and skull base osteomyelitis. Therefore, active treatment based on early diagnosis is essential. This study developed a multi-modal multi-fusion (MMMF) model that automatically diagnoses ear diseases by applying endoscopic images of the tympanic membrane (TM) and pure-tone audiometry (PTA) data to a deep learning model. The primary aim of the proposed MMMF model is a

OtorhinolaryngologyMedicine
2
Article|7 citations·2023
Multiclass Classification by Various Machine Learning Algorithms and Interpretation of the Risk Factors of Pedestrian Accidents Using Explainable AI
Sanghun Lee, Sangyeop Kim, Jaehoon Kim, Doyun Kim, Dohyun Lee, Gwangmuk Im, Hyeonseop Yuk, Tae‐Young Heo
SJR Q2Mathematical Problems in EngineeringOA

Pedestrian injuries and fatalities due to traffic accidents remain at a high level. Therefore, the need for efforts to reduce this ratio is on the rise. Machine learning models can facilitate the exploration of the various factors that influence the occurrence of pedestrian accidents. In this study, we used data on pedestrian traffic accidents classified into three categories of injury severity: minor, severe, and fatal. To compare the performance of various types of models, logistic regression,

Safety, Risk, Reliability and QualityEngineering
3
Article|3 citations·2025
Human-guided collective LLM intelligence for strategic planning via two-stage information retrieval
Sangyeop Kim, Jinxuan Ha, Hangyeul Lee, Sohhyung Park, Sungzoon Cho
SJR Q1Information Processing & Management
Management Information SystemsBusiness, Management and Accounting
4
Article|2 citations·2022
An alternative testing method to investigate creep-dominant creep-fatigue interaction and its application on modified 9Cr-1Mo steel
Uijeong Ro, Jeong Hwan Kim, Sangyeop Kim, Moon Ki Kim
SJR Q2Journal of Mechanical Science and Technology
Mechanical EngineeringEngineering
5
Article|1 citations·2025
LLM-guided Plan and Retrieval: A Strategic Alignment for Interpretable User Satisfaction Estimation in Dialogue
Sangyeop Kim, Sohhyung Park, Jaewon Jung, Jinseok Kim, Sungzoon Cho
OA

Sangyeop Kim, Sohhyung Park, Jaewon Jung, Jinseok Kim, Sungzoon Cho. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.

Artificial IntelligenceComputer Science
6
Preprint|1 citations·2023
Automatic Diagnosis of Chronic Otitis Media with a Dual Neural Network Using Pure-Tone Audiometry and Tympanic Membrane Images
Tae‐Wan Kim, Sangyeop Kim, Jaeyoung Kim, Yeonjoon Lee, June Choi
SSRN Electronic JournalOA
OtorhinolaryngologyMedicine
7
Article|1 citations·2025
HEISIR: Hierarchical Expansion of Inverted Semantic Indexing for Training-free Retrieval of Conversational Data using LLMs
Sangyeop Kim, H. P. Lee, Yohan Lee
OA

The growth of conversational AI services has increased demand for effective information retrieval from dialogue data.However, existing methods often face challenges in capturing semantic intent or require extensive labeling and fine-tuning.This paper introduces HEISIR (Hierarchical Expansion of Inverted Semantic Indexing for Retrieval), a novel framework that enhances semantic understanding in conversational data retrieval through optimized data ingestion, eliminating the need for resource-inten

Artificial IntelligenceComputer Science
8
Preprint|0 citations·2025
What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
Sangyeop Kim, Yo Han Lee, Yun‐Heub Song, Kimin Lee
ArXiv.orgOA

We investigate long-context vulnerabilities in Large Language Models (LLMs) through Many-Shot Jailbreaking (MSJ). Our experiments utilize context length of up to 128K tokens. Through comprehensive analysis with various many-shot attack settings with different instruction styles, shot density, topic, and format, we reveal that context length is the primary factor determining attack effectiveness. Critically, we find that successful attacks do not require carefully crafted harmful content. Even re

Information SystemsComputer Science
9
Preprint|0 citations·2025
HEISIR: Hierarchical Expansion of Inverted Semantic Indexing for Training-free Retrieval of Conversational Data using LLMs
Sangyeop Kim, H. P. Lee, Yohan Lee
ArXiv.orgOA

The growth of conversational AI services has increased demand for effective information retrieval from dialogue data. However, existing methods often face challenges in capturing semantic intent or require extensive labeling and fine-tuning. This paper introduces HEISIR (Hierarchical Expansion of Inverted Semantic Indexing for Retrieval), a novel framework that enhances semantic understanding in conversational data retrieval through optimized data ingestion, eliminating the need for resource-int

Artificial IntelligenceComputer Science
10
Article|0 citations·2025
Pre-Storage Reasoning for Episodic Memory: Shifting Inference Burden to Memory for Personalized Dialogue
Sangyeop Kim, Yo Han Lee, Sang-Hwa Kim, Hyun-Jong Kim, Sungzoon Cho
OA

Effective long-term memory in conversational AI requires synthesizing information across multiple sessions.However, current systems place excessive reasoning burden on response generation, making performance significantly dependent on model sizes.We introduce PRE-Mem (Pre-storage Reasoning for Episodic Memory), a novel approach that shifts complex reasoning processes from inference to memory construction.PREMem extracts finegrained memory fragments categorized into factual, experiential, and sub

Artificial IntelligenceComputer Science
11
Article|0 citations·2024
Safe-Embed: Unveiling the Safety-Critical Knowledge of Sentence Encoders
Jinseok Kim, Jaewon Jung, Sangyeop Kim, Sohhyung Park, Sungzoon Cho
OA
Artificial IntelligenceComputer Science
12
Article|0 citations·2025
What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
Sangyeop Kim, Yo Han Lee, Yun‐Heub Song, Kimin Lee
OA

We investigate long-context vulnerabilities in Large Language Models (LLMs) through Many-Shot Jailbreaking (MSJ).Our experiments utilize context length of up to 128K tokens.Through comprehensive analysis with various many-shot attack settings with different instruction styles, shot density, topic, and format, we reveal that context length is the primary factor determining attack effectiveness.Critically, we find that successful attacks do not require carefully crafted harmful content.Even repeti

Information SystemsComputer Science
13
Preprint|0 citations·2025
Finding Diamonds in Conversation Haystacks: A Benchmark for Conversational Data Retrieval
Yohan Lee, Song Yongping, Sangyeop Kim
arXiv (Cornell University)OA

We present the Conversational Data Retrieval (CDR) benchmark, the first comprehensive test set for evaluating systems that retrieve conversation data for product insights. With 1.6k queries across five analytical tasks and 9.1k conversations, our benchmark provides a reliable standard for measuring conversational data retrieval performance. Our evaluation of 16 popular embedding models shows that even the best models reach only around NDCG@10 of 0.51, revealing a substantial gap between document

Artificial IntelligenceComputer Science
14
Preprint|0 citations·2025
LLM-guided Plan and Retrieval: A Strategic Alignment for Interpretable User Satisfaction Estimation in Dialogue
Sangyeop Kim, Sohhyung Park, Jaewon Jung, Jinseok Kim, Sungzoon Cho
ArXiv.orgOA

Understanding user satisfaction with conversational systems, known as User Satisfaction Estimation (USE), is essential for assessing dialogue quality and enhancing user experiences. However, existing methods for USE face challenges due to limited understanding of underlying reasons for user dissatisfaction and the high costs of annotating user intentions. To address these challenges, we propose PRAISE (Plan and Retrieval Alignment for Interpretable Satisfaction Estimation), an interpretable fram

Artificial IntelligenceComputer Science
15
Preprint|0 citations·2025
Understanding User Perception of Human-Llm Collaboration in Ai-Assisted Decision-Making
Sohhyung Park, Jongwon Ha, Hangyeul Lee, Sangyeop Kim, Sungzoon Cho
SSRN Electronic JournalOA
Social PsychologyPsychology

Research Areas

Artificial IntelligenceComputer Vision and Pattern RecognitionInformation SystemsOtorhinolaryngologySafety, Risk, Reliability and QualityManagement Information Systems

Sang-Young Kimの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。