Skip to main content

Seunghoon Hong

Korea Advanced Institute of Science and Technology · 情報科学

研究室紹介

Professor Seunghoon Hong's research lab specializes in probabilistic modeling and generative AI, with a focus on improving the efficiency, fidelity, and robustness of diffusion and flow-based generative models. The lab explores fundamental challenges in latent space modeling, such as distribution shift in variable-length tokenization, attention dynamics in vision-language models, and geometric constraints in transformation inversion on Lie groups. By integrating insights from differential geometry, information theory, and deep learning, the lab develops training-free and geometrically principled methods that enhance sample quality and representational alignment.

diffusion modelsflow-based generative modelsLie group diffusionlatent variable modelingattention sinks

Research Overview

Papers
8
Total Citations
0
Papers (5y)
8
Primary Field
情報科学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
8total
2026
Citations per year (5y)
0total
2026

Selected Papers

8
1
Article|0 citations·2026
Training-Free Refinement of Flow Matching with Divergence-based Sampling
Yeonwoo Cha, Jaehoon Yoo, Semin Kim, Yunseo Park, Jinhyeon Kwon, Seunghoon Hong
arXiv (Cornell University)OA

Flow-based models learn a target distribution by modeling a marginal velocity field, defined as the average of sample-wise velocities connecting each sample from a simple prior to the target data. When sample-wise velocities conflict at the same intermediate state, however, this averaged velocity can misguide samples toward low-density regions, degrading generation quality. To address this issue, we propose the Flow Divergence Sampler (FDS), a training-free framework that refines intermediate st

Computer Vision and Pattern RecognitionComputer Science
2
Preprint|0 citations·2026
Training-Free Refinement of Flow Matching with Divergence-based Sampling
Yeonwoo Cha, Jaehoon Yoo, Kim, Semin, Yunseo Park, Jinhyeon Kwon, Seunghoon Hong
arXiv (Cornell University)OA

Flow-based models learn a target distribution by modeling a marginal velocity field, defined as the average of sample-wise velocities connecting each sample from a simple prior to the target data. When sample-wise velocities conflict at the same intermediate state, however, this averaged velocity can misguide samples toward low-density regions, degrading generation quality. To address this issue, we propose the Flow Divergence Sampler (FDS), a training-free framework that refines intermediate st

Computer Vision and Pattern RecognitionComputer Science
3
Preprint|0 citations·2026
Variable-Length Tokenization via Learnable Global Merging for Diffusion Transformers
D Lee, Seunghoon Hong
arXiv (Cornell University)OA

Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality-compute trade-off is largely constrained by the tokenizer's fixed compression ratio. Variable-length tokenizers (VLTs) promise adaptive compression by varying token counts, allowing diffusion models to flexibly balance quality and compute. However, conventional VLTs modulate length by truncating ordered token sequences, which makes token semantics depend on token position and breaks representational alignm

Computer Vision and Pattern RecognitionComputer Science
4
Article|0 citations·2026
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
Jiho Choi, Jaemin Kim, Sanghwan Kim, Seunghoon Hong, Jin-Hwi Park
arXiv (Cornell University)OA

Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their cross-modal impact in Large Vision-Language Models (LVLM) remains largely unexplored: are they redundant artifacts or essential global priors? This paper first categorizes visual sinks into two distinct categories: ViT-emerged sinks (V-sinks), which propagate from the vision encoder, and LLM-emerged sinks (L-sinks), which arise within deep LLM layers

Computer Vision and Pattern RecognitionComputer Science
5
Preprint|0 citations·2026
Inverting Data Transformations via Diffusion Sampling
Jinwoo Kim, Sékou-Oumar Kaba, Jiyun Park, Seunghoon Hong, Siamak Ravanbakhsh
SJR Q1Open MINDOA

We study the problem of transformation inversion on general Lie groups: a datum is transformed by an unknown group element, and the goal is to recover an inverse transformation that maps it back to the original data distribution. Such unknown transformations arise widely in machine learning and scientific modeling, where they can significantly distort observations. We take a probabilistic view and model the posterior over transformations as a Boltzmann distribution defined by an energy function

Computer Vision and Pattern RecognitionComputer Science
6
Article|0 citations·2026
Inverting Data Transformations via Diffusion Sampling
Jinwoo Kim, Sékou-Oumar Kaba, Jiyun Park, Seunghoon Hong, Siamak Ravanbakhsh
ArXiv.orgOA

We study the problem of transformation inversion on general Lie groups: a datum is transformed by an unknown group element, and the goal is to recover an inverse transformation that maps it back to the original data distribution. Such unknown transformations arise widely in machine learning and scientific modeling, where they can significantly distort observations. We take a probabilistic view and model the posterior over transformations as a Boltzmann distribution defined by an energy function

Computer Vision and Pattern RecognitionComputer Science
7
Preprint|0 citations·2026
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
Jiho Choi, Jaemin Kim, Sanghwan Kim, Seunghoon Hong, Jin-Hwi Park
arXiv (Cornell University)OA

Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their cross-modal impact in Large Vision-Language Models (LVLM) remains largely unexplored: are they redundant artifacts or essential global priors? This paper first categorizes visual sinks into two distinct categories: ViT-emerged sinks (V-sinks), which propagate from the vision encoder, and LLM-emerged sinks (L-sinks), which arise within deep LLM layers

Computer Vision and Pattern RecognitionComputer Science
8
Article|0 citations·2026
Infinite Mask Diffusion for Few-Step Distillation
Jaehoon Yoo, Wonjung Kim, Chanhyuk Lee, Seunghoon Hong
ArXiv.orgOA

Masked Diffusion Models (MDMs) have emerged as a promising alternative to autoregressive models in language modeling, offering the advantages of parallel decoding and bidirectional context processing within a simple yet effective framework. Specifically, their explicit distinction between masked tokens and data underlies their simple framework and effective conditional generation. However, MDMs typically require many sampling iterations due to factorization errors stemming from simultaneous toke

Artificial IntelligenceComputer Science

Research Areas

Computer Vision and Pattern RecognitionArtificial Intelligence

Seunghoon Hongの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。