Skip to main content

Hankook Lee

Sungkyunkwan University · 情報科学

研究室紹介

Professor Hankook Lee's research lab specializes in advancing machine learning for robust and trustworthy AI, with a strong focus on anomaly detection, out-of-distribution generalization, and trustworthy generation in vision and language. The lab develops innovative frameworks that integrate contrastive representation learning, diffusion models, and retrieval-augmented generation to improve model reliability in real-world scenarios where data is scarce or biased. Key research directions include semantic-aware outlier generation, context-aware tabular anomaly detection, and enhancing large language models through external knowledge grounding.

anomaly detectionout-of-distribution detectiondiffusion modelsretrieval-augmented generationcontrastive representation learning

Research Overview

Papers
11
Total Citations
12
Papers (5y)
11
Primary Field
情報科学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
11total
2023
2024
2025
Citations per year (5y)
12total
202320242025

Selected Papers

11
1
Article|6 citations·2024
Few-Shot Anomaly Detection via Personalization
Sangkyung Kwak, Jongheon Jeong, Hankook Lee, W. Kim, Dongho Seo, Woo-Seong Yun, Won Jin Lee, Jinwoo Shin
SJR Q1IEEE AccessOA

Even with a plenty amount of normal samples, anomaly detection has been considered as a challenging machine learning task due to its one-class nature, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i.e</i> ., the lack of anomalous samples in training time. It is only recently that a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">few-shot</i> regime of anomaly detection became feasible in this re

Artificial IntelligenceComputer Science
2
Preprint|4 citations·2023
Guiding Energy-based Models via Contrastive Latent Variables
Hankook Lee, Jongheon Jeong, Sejun Park, Jinwoo Shin
arXiv (Cornell University)OA

An energy-based model (EBM) is a popular generative framework that offers both explicit density and architectural flexibility, but training them is difficult since it is often unstable and time-consuming. In recent years, various training techniques have been developed, e.g., better divergence measures or stabilization in MCMC sampling, but there often exists a large gap between EBMs and other generative frameworks like GANs in terms of generation quality. In this paper, we propose a novel and e

Computer Vision and Pattern RecognitionComputer Science
3
Article|1 citations·2025
Rescuing the Unpoisoned: Efficient Defense Against Knowledge Corruption Attacks on RAG Systems
Minseok Kim, Hankook Lee, Hyungjoon Koo

Large language models (LLMs) are reshaping numerous facets of our daily lives, leading to their widespread adoption as web-based services. Despite their versatility, LLMs face notable challenges, such as generating hallucinated content and lacking access to up-to-date information. Lately, to address such limitations, Retrieval-Augmented Generation (RAG) has emerged as a promising direction by generating responses grounded in external knowledge sources. A typical RAG system consists of i) a retri

Artificial IntelligenceComputer Science
4
Article|1 citations·2025
Diffusion-based Semantic Outlier Generation via Nuisance Awareness for Out-of-Distribution Detection
Suhee Yoon, Sanghyu Yoon, Ye Seul Sim, Sungik Choi, Kyungeun Lee, Hye-Seung Cho, Hankook Lee, Woohyung Lim
Proceedings of the AAAI Conference on Artificial IntelligenceOA

Out-of-distribution (OOD) detection, determining whether a given sample is part of the in-distribution (ID) or not, has been newly explored by a generative model-based outlier synthesizing approach, especially with diffusion models. Nonetheless, existing diffusion models often produce outliers that are considerably distant from the ID in pixel-space, showing limited efficacy for capturing subtle distinctions between ID and OOD. To address these issues, we propose a novel framework, Semantic Outl

Artificial IntelligenceComputer Science
5
Preprint|0 citations·2025
TimePerceiver: An Encoder-Decoder Framework for Generalized Time-Series Forecasting
Jaebin Lee, Hankook Lee
arXiv (Cornell University)OA

In machine learning, effective modeling requires a holistic consideration of how to encode inputs, make predictions (i.e., decoding), and train the model. However, in time-series forecasting, prior work has predominantly focused on encoder design, often treating prediction and training as separate or secondary concerns. In this paper, we propose TimePerceiver, a unified encoder-decoder forecasting framework that is tightly aligned with an effective training strategy. To be specific, we first gen

Signal ProcessingComputer Science
6
Preprint|0 citations·2025
ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection
Sanghyu Yoon, Dongmin Kim, Suhee Yoon, Ye Seul Sim, Seungdong Yoa, Hye-Seung Cho, Soon Young Lee, Hankook Lee, Woohyung Lim
ArXiv.orgOA

In tabular anomaly detection (AD), textual semantics often carry critical signals, as the definition of an anomaly is closely tied to domain-specific context. However, existing benchmarks provide only raw data points without semantic context, overlooking rich textual metadata such as feature descriptions and domain knowledge that experts rely on in practice. This limitation restricts research flexibility and prevents models from fully leveraging domain knowledge for detection. ReTabAD addresses

Artificial IntelligenceComputer Science
7
Article|0 citations·2023
Enhancing Multiple Reliability Measures via Nuisance-Extended Information Bottleneck
Jongheon Jeong, Sihyun Yu, Hankook Lee, Jinwoo Shin

In practical scenarios where training data is limited, many predictive signals in the data can be rather from some biases in data acquisition (i.e., less generalizable), so that one cannot prevent a model from co-adapting on such (so-called) “shortcut” signals: this makes the model fragile in various distribution shifts. To bypass such failure modes, we consider an adversarial threat model under a mutual information constraint to cover a wider class of perturbations in training. This motivates u

Artificial IntelligenceComputer Science
8
Article|0 citations·2025
TimePerceiver: An Encoder-Decoder Framework for Generalized Time-Series Forecasting
Jaebin Lee, Hankook Lee
ArXiv.orgOA

In machine learning, effective modeling requires a holistic consideration of how to encode inputs, make predictions (i.e., decoding), and train the model. However, in time-series forecasting, prior work has predominantly focused on encoder design, often treating prediction and training as separate or secondary concerns. In this paper, we propose TimePerceiver, a unified encoder-decoder forecasting framework that is tightly aligned with an effective training strategy. To be specific, we first gen

Signal ProcessingComputer Science
9
Preprint|0 citations·2023
Enhancing Multiple Reliability Measures via Nuisance-extended Information Bottleneck
Jongheon Jeong, Sihyun Yu, Hankook Lee, Jinwoo Shin
arXiv (Cornell University)OA

In practical scenarios where training data is limited, many predictive signals in the data can be rather from some biases in data acquisition (i.e., less generalizable), so that one cannot prevent a model from co-adapting on such (so-called) "shortcut" signals: this makes the model fragile in various distribution shifts. To bypass such failure modes, we consider an adversarial threat model under a mutual information constraint to cover a wider class of perturbations in training. This motivates u

Artificial IntelligenceComputer Science
10
Preprint|0 citations·2024
Partial-Multivariate Model for Forecasting
Jaehoon Lee, Hankook Lee, Sunho Choi, Sungjun Cho, Minjoo Lee
arXiv (Cornell University)OA

When solving forecasting problems including multiple time-series features, existing approaches often fall into two extreme categories, depending on whether to utilize inter-feature information: univariate and complete-multivariate models. Unlike univariate cases which ignore the information, complete-multivariate models compute relationships among a complete set of features. However, despite the potential advantage of leveraging the additional information, complete-multivariate models sometimes

Management Science and Operations ResearchDecision Sciences
11
Preprint|0 citations·2024
Diffusion based Semantic Outlier Generation via Nuisance Awareness for Out-of-Distribution Detection
Suhee Yoon, Sanghyu Yoon, Ye Seul Sim, Sungik Choi, Kyung‐Eun Lee, Hye-Seung Cho, Hankook Lee, Woohyung Lim
arXiv (Cornell University)OA

Out-of-distribution (OOD) detection, which determines whether a given sample is part of the in-distribution (ID), has recently shown promising results through training with synthetic OOD datasets. Nonetheless, existing methods often produce outliers that are considerably distant from the ID, showing limited efficacy for capturing subtle distinctions between ID and OOD. To address these issues, we propose a novel framework, Semantic Outlier generation via Nuisance Awareness (SONA), which notably

Artificial IntelligenceComputer Science

Research Areas

Artificial IntelligenceSignal ProcessingComputer Vision and Pattern RecognitionManagement Science and Operations Research

Hankook Leeの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。