Sang-Young Kim
Ewha Womans University · 情報科学
研究室紹介
Professor Sang-Young Kim's research lab specializes in advancing artificial intelligence with a focus on multimodal learning, conversational AI, and the safety and robustness of large language models. The lab develops innovative deep learning and machine learning frameworks that enhance diagnostic accuracy in healthcare through multimodal data fusion, such as combining medical imaging and audiometric data. It also pioneers efficient, low-resource methods for semantic retrieval in dialogue systems, exemplified by the HEISIR framework, while critically investigating vulnerabilities in long-context LLMs to improve model safety. The lab's work bridges practical applications in healthcare and natural language processing with foundational research in AI reliability and interpretability.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15Chronic otitis media is characterized by recurrent infections, leading to serious complications, such as meningitis, facial palsy, and skull base osteomyelitis. Therefore, active treatment based on early diagnosis is essential. This study developed a multi-modal multi-fusion (MMMF) model that automatically diagnoses ear diseases by applying endoscopic images of the tympanic membrane (TM) and pure-tone audiometry (PTA) data to a deep learning model. The primary aim of the proposed MMMF model is a
Pedestrian injuries and fatalities due to traffic accidents remain at a high level. Therefore, the need for efforts to reduce this ratio is on the rise. Machine learning models can facilitate the exploration of the various factors that influence the occurrence of pedestrian accidents. In this study, we used data on pedestrian traffic accidents classified into three categories of injury severity: minor, severe, and fatal. To compare the performance of various types of models, logistic regression,
Sangyeop Kim, Sohhyung Park, Jaewon Jung, Jinseok Kim, Sungzoon Cho. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
The growth of conversational AI services has increased demand for effective information retrieval from dialogue data.However, existing methods often face challenges in capturing semantic intent or require extensive labeling and fine-tuning.This paper introduces HEISIR (Hierarchical Expansion of Inverted Semantic Indexing for Retrieval), a novel framework that enhances semantic understanding in conversational data retrieval through optimized data ingestion, eliminating the need for resource-inten
We investigate long-context vulnerabilities in Large Language Models (LLMs) through Many-Shot Jailbreaking (MSJ). Our experiments utilize context length of up to 128K tokens. Through comprehensive analysis with various many-shot attack settings with different instruction styles, shot density, topic, and format, we reveal that context length is the primary factor determining attack effectiveness. Critically, we find that successful attacks do not require carefully crafted harmful content. Even re
The growth of conversational AI services has increased demand for effective information retrieval from dialogue data. However, existing methods often face challenges in capturing semantic intent or require extensive labeling and fine-tuning. This paper introduces HEISIR (Hierarchical Expansion of Inverted Semantic Indexing for Retrieval), a novel framework that enhances semantic understanding in conversational data retrieval through optimized data ingestion, eliminating the need for resource-int
Effective long-term memory in conversational AI requires synthesizing information across multiple sessions.However, current systems place excessive reasoning burden on response generation, making performance significantly dependent on model sizes.We introduce PRE-Mem (Pre-storage Reasoning for Episodic Memory), a novel approach that shifts complex reasoning processes from inference to memory construction.PREMem extracts finegrained memory fragments categorized into factual, experiential, and sub
We investigate long-context vulnerabilities in Large Language Models (LLMs) through Many-Shot Jailbreaking (MSJ).Our experiments utilize context length of up to 128K tokens.Through comprehensive analysis with various many-shot attack settings with different instruction styles, shot density, topic, and format, we reveal that context length is the primary factor determining attack effectiveness.Critically, we find that successful attacks do not require carefully crafted harmful content.Even repeti
We present the Conversational Data Retrieval (CDR) benchmark, the first comprehensive test set for evaluating systems that retrieve conversation data for product insights. With 1.6k queries across five analytical tasks and 9.1k conversations, our benchmark provides a reliable standard for measuring conversational data retrieval performance. Our evaluation of 16 popular embedding models shows that even the best models reach only around NDCG@10 of 0.51, revealing a substantial gap between document
Understanding user satisfaction with conversational systems, known as User Satisfaction Estimation (USE), is essential for assessing dialogue quality and enhancing user experiences. However, existing methods for USE face challenges due to limited understanding of underlying reasons for user dissatisfaction and the high costs of annotating user intentions. To address these challenges, we propose PRAISE (Plan and Retrieval Alignment for Interpretable Satisfaction Estimation), an interpretable fram