Skip to main content

홍승훈 교수

Seunghoon Hong

KAIST 김재철AI대학원 · 컴퓨터과학

연구실 소개

홍승훈 교수의 연구실은 확률적 추론과 미분기하학을 기반으로 한 새로운 생성 모델링 기법을 개발하고 있습니다. 특히 흐름 기반 모델의 중간 상태에서의 속도 갈등 문제를 해결하기 위한 훈련 없이도 성능을 향상시키는 '플로우 발산 샘플러'(FDS)를 비롯해, 리만 다양체 위에서의 확산 프로세스를 활용한 변환 복원 기법을 연구하고 있습니다. 또한 뷰어-언어 모델에서의 주목력 집중 현상과 변수 길이 토크나이저를 통한 품질-계산량 트레이드오��� 최적화 등, 다중 모odal 및 효율적 생성 아키텍처의 이론적 기반을 탐구하고 있습니다. 이 모든 연구는 데이터의 기하학적 구조를 정확히 이해하고 활용하는 데 초점을 맞추고 있습니다.

흐름 기반 모델리만 다양체확산 프로세스변환 복원변수 길이 토크나이저

연구 현황

논문 수
8
총 인용 수
0
최근 5년 논문
8
주요 분야
컴퓨터과학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
8총합
2026
5개년 연도별 피인용 수
0총합
2026

주요 논문

8
1
논문|인용수 0·2026
Training-Free Refinement of Flow Matching with Divergence-based Sampling
Yeonwoo Cha, Jaehoon Yoo, Semin Kim, Yunseo Park, Jinhyeon Kwon, Seunghoon Hong
arXiv (Cornell University)OA

Flow-based models learn a target distribution by modeling a marginal velocity field, defined as the average of sample-wise velocities connecting each sample from a simple prior to the target data. When sample-wise velocities conflict at the same intermediate state, however, this averaged velocity can misguide samples toward low-density regions, degrading generation quality. To address this issue, we propose the Flow Divergence Sampler (FDS), a training-free framework that refines intermediate st

Computer Vision and Pattern RecognitionComputer Science
2
preprint|인용수 0·2026
Training-Free Refinement of Flow Matching with Divergence-based Sampling
Yeonwoo Cha, Jaehoon Yoo, Kim, Semin, Yunseo Park, Jinhyeon Kwon, Seunghoon Hong
arXiv (Cornell University)OA

Flow-based models learn a target distribution by modeling a marginal velocity field, defined as the average of sample-wise velocities connecting each sample from a simple prior to the target data. When sample-wise velocities conflict at the same intermediate state, however, this averaged velocity can misguide samples toward low-density regions, degrading generation quality. To address this issue, we propose the Flow Divergence Sampler (FDS), a training-free framework that refines intermediate st

Computer Vision and Pattern RecognitionComputer Science
3
preprint|인용수 0·2026
Variable-Length Tokenization via Learnable Global Merging for Diffusion Transformers
D Lee, Seunghoon Hong
arXiv (Cornell University)OA

Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality-compute trade-off is largely constrained by the tokenizer's fixed compression ratio. Variable-length tokenizers (VLTs) promise adaptive compression by varying token counts, allowing diffusion models to flexibly balance quality and compute. However, conventional VLTs modulate length by truncating ordered token sequences, which makes token semantics depend on token position and breaks representational alignm

Computer Vision and Pattern RecognitionComputer Science
4
논문|인용수 0·2026
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
Jiho Choi, Jaemin Kim, Sanghwan Kim, Seunghoon Hong, Jin-Hwi Park
arXiv (Cornell University)OA

Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their cross-modal impact in Large Vision-Language Models (LVLM) remains largely unexplored: are they redundant artifacts or essential global priors? This paper first categorizes visual sinks into two distinct categories: ViT-emerged sinks (V-sinks), which propagate from the vision encoder, and LLM-emerged sinks (L-sinks), which arise within deep LLM layers

Computer Vision and Pattern RecognitionComputer Science
5
preprint|인용수 0·2026
Inverting Data Transformations via Diffusion Sampling
Jinwoo Kim, Sékou-Oumar Kaba, Jiyun Park, Seunghoon Hong, Siamak Ravanbakhsh
SJR Q1Open MINDOA

We study the problem of transformation inversion on general Lie groups: a datum is transformed by an unknown group element, and the goal is to recover an inverse transformation that maps it back to the original data distribution. Such unknown transformations arise widely in machine learning and scientific modeling, where they can significantly distort observations. We take a probabilistic view and model the posterior over transformations as a Boltzmann distribution defined by an energy function

Computer Vision and Pattern RecognitionComputer Science
6
논문|인용수 0·2026
Inverting Data Transformations via Diffusion Sampling
Jinwoo Kim, Sékou-Oumar Kaba, Jiyun Park, Seunghoon Hong, Siamak Ravanbakhsh
ArXiv.orgOA

We study the problem of transformation inversion on general Lie groups: a datum is transformed by an unknown group element, and the goal is to recover an inverse transformation that maps it back to the original data distribution. Such unknown transformations arise widely in machine learning and scientific modeling, where they can significantly distort observations. We take a probabilistic view and model the posterior over transformations as a Boltzmann distribution defined by an energy function

Computer Vision and Pattern RecognitionComputer Science
7
preprint|인용수 0·2026
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
Jiho Choi, Jaemin Kim, Sanghwan Kim, Seunghoon Hong, Jin-Hwi Park
arXiv (Cornell University)OA

Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their cross-modal impact in Large Vision-Language Models (LVLM) remains largely unexplored: are they redundant artifacts or essential global priors? This paper first categorizes visual sinks into two distinct categories: ViT-emerged sinks (V-sinks), which propagate from the vision encoder, and LLM-emerged sinks (L-sinks), which arise within deep LLM layers

Computer Vision and Pattern RecognitionComputer Science
8
논문|인용수 0·2026
Infinite Mask Diffusion for Few-Step Distillation
Jaehoon Yoo, Wonjung Kim, Chanhyuk Lee, Seunghoon Hong
ArXiv.orgOA

Masked Diffusion Models (MDMs) have emerged as a promising alternative to autoregressive models in language modeling, offering the advantages of parallel decoding and bidirectional context processing within a simple yet effective framework. Specifically, their explicit distinction between masked tokens and data underlies their simple framework and effective conditional generation. However, MDMs typically require many sampling iterations due to factorization errors stemming from simultaneous toke

Artificial IntelligenceComputer Science

대표 연구 분야

Computer Vision and Pattern RecognitionArtificial Intelligence

홍승훈 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.