Skip to main content

한승열 교수

Seungyul Han

UNIST 전기전자공학과 · 컴퓨터과학

연구실 소개

한승열 교수의 연구실은 강화학습 기반의 최적화 기법을 핵심으로 하며, 특히 다중 에이gent 환경에서의 탐색 효율성 향상과 오프-폴리시 학습의 안정성 확보에 초점을 맞추고 있습니다. GFDM 신호 처리에서의 필터 최적화부터, 딥 강화학습을 활용한 큐 안정성 제어, 다중 에이gent 환경에서의 형성 기반 탐색 기법까지 응용 분야가 다양합니다. 특히 정책 엔트로피 정규화, 중요도 샘플링 가중치의 차원별 클리핑, 경험 리플레이의 다중 배치 설계 등 학습 효율성과 수렴 안정성을 높이는 기초 알고리즘 개선에 뛰어난 기여를 하고 있습니다.

강화학습 최적화다중 에이gent 탐색오프-폴리시 학습정책 엔트로피경험 리플레이

연구 현황

논문 수
40
총 인용 수
116
최근 5년 논문
30
주요 분야
컴퓨터과학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
30총합
2022
2023
2024
2025
2026
5개년 연도별 피인용 수
12총합
20222023202420252026

주요 논문

15
1
논문|인용수 60·2016
Filter Design for Generalized Frequency-Division Multiplexing
Seungyul Han, Youngchul Sung, Yong H. Lee
SJR Q1IEEE Transactions on Signal ProcessingOA

In this paper, optimal filter design for generalized frequency-division multiplexing (GFDM) is considered under two design criteria: rate maximization and out-of-band (OOB) emission minimization. First, the problem of GFDM filter optimization for rate maximization is formulated by expressing the transmission rate of GFDM as a function of GFDM filter coefficients. It is shown that Dirichlet filters are rate-optimal in additive white Gaussian noise channels with no carrier frequency offset (CFO) u

Electrical and Electronic EngineeringEngineering
2
preprint|인용수 10·2020
Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration
Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

In this paper, sample-aware policy entropy regularization is proposed to enhance the conventional policy entropy regularization for better exploration. Exploiting the sample distribution obtainable from the replay buffer, the proposed sample-aware entropy regularization maximizes the entropy of the weighted sum of the policy action distribution and the sample action distribution from the replay buffer for sample-efficient exploration. A practical algorithm named diversity actor-critic (DAC) is d

Artificial IntelligenceComputer Science
3
preprint|인용수 9·2020
A Reinforcement Learning Formulation of the Lyapunov Optimization: Application to Edge Computing Systems with Queue Stability
Sohee Bae, Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

In this paper, a deep reinforcement learning (DRL)-based approach to the Lyapunov optimization is considered to minimize the time-average penalty while maintaining queue stability. A proper construction of state and action spaces is provided to form a proper Markov decision process (MDP) for the Lyapunov optimization. A condition for the reward function of reinforcement learning (RL) for queue stability is derived. Based on the analysis and practical RL with reward discounting, a class of reward

Computer Networks and CommunicationsComputer Science
4
논문|인용수 7·2024
FoX: Formation-Aware Exploration in Multi-Agent Reinforcement Learning
Yonghyeon Jo, Sunwoo Lee, Junghyuk Yeom, Seungyul Han
Proceedings of the AAAI Conference on Artificial IntelligenceOA

Recently, deep multi-agent reinforcement learning (MARL) has gained significant popularity due to its success in various cooperative multi-agent tasks. However, exploration still remains a challenging problem in MARL due to the partial observability of the agents and the exploration space that can grow exponentially as the number of agents increases. Firstly, in order to address the scalability issue of the exploration space, we define a formation-based equivalence relation on the exploration sp

Artificial IntelligenceComputer Science
5
논문|인용수 6·2019
Dimension-Wise Importance Sampling Weight Clipping for Sample-Efficient Reinforcement Learning
Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

In importance sampling (IS)-based reinforcement learning algorithms such as Proximal Policy Optimization (PPO), IS weights are typically clipped to avoid large variance in learning. However, policy update from clipped statistics induces large bias in tasks with high action dimensions, and bias from clipping makes it difficult to reuse old samples with large IS weights. In this paper, we consider PPO, a representative on-policy algorithm, and propose its improvement by dimension-wise IS weight cl

Artificial IntelligenceComputer Science
6
preprint|인용수 5·2017
Multi-Batch Experience Replay for Fast Convergence of Continuous Action Control.
Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

Policy gradient methods for direct policy optimization are widely considered to obtain optimal policies in continuous Markov decision process (MDP) environments. However, policy gradient methods require exponentially many samples as the dimension of the action space increases. Thus, off-policy learning with experience replay is proposed to enable the agent to learn by using samples of other policies. Generally, large replay memories are preferred to minimize the sample correlation but large repl

Artificial IntelligenceComputer Science
7
preprint|인용수 5·2017
AMBER: Adaptive Multi-Batch Experience Replay for Continuous Action Control
Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

In this paper, a new adaptive multi-batch experience replay scheme is proposed for proximal policy optimization (PPO) for continuous action control. On the contrary to original PPO, the proposed scheme uses the batch samples of past policies as well as the current policy for the update for the next policy, where the number of the used past batches is adaptively determined based on the oldness of the past batches measured by the average importance sampling (IS) weight. The new algorithm construct

Artificial IntelligenceComputer Science
8
preprint|인용수 5·2021
A Max-Min Entropy Framework for Reinforcement Learning
Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

In this paper, we propose a max-min entropy framework for reinforcement learning (RL) to overcome the limitation of the soft actor-critic (SAC) algorithm implementing the maximum entropy RL in model-free sample-based learning. Whereas the maximum entropy RL guides learning for policies to reach states with high entropy in the future, the proposed max-min entropy framework aims to learn to visit states with low entropy and maximize the entropy of these low-entropy states to promote better explora

Artificial IntelligenceComputer Science
9
preprint|인용수 3·2022
Robust Imitation Learning against Variations in Environment Dynamics
Jong-Seong Chae, Seungyul Han, Whiyoung Jung, Myungsik Cho, Sungho Choi, Youngchul Sung
arXiv (Cornell University)OA

In this paper, we propose a robust imitation learning (IL) framework that improves the robustness of IL when environment dynamics are perturbed. The existing IL framework trained in a single environment can catastrophically fail with perturbations in environment dynamics because it does not capture the situation that underlying environment dynamics can be changed. Our framework effectively deals with environments with varying dynamics by imitating multiple experts in sampled environment dynamics

Artificial IntelligenceComputer Science
10
preprint|인용수 3·2020
Cross-Domain Imitation Learning with a Dual Structure
Sungho Choi, Seungyul Han, Woojun Kim, Youngchul Sung
arXiv (Cornell University)OA

In this paper, we consider cross-domain imitation learning (CDIL) in which an agent in a target domain learns a policy to perform well in the target domain by observing expert demonstrations in a source domain without accessing any reward function. In order to overcome the domain difference for imitation learning, we propose a dual-structured learning method. The proposed learning method extracts two feature vectors from each input observation such that one vector contains domain information and

Artificial IntelligenceComputer Science
11
논문|인용수 2·2025
Adaptive multi-model fusion learning for sparse-reward reinforcement learning
Giseung Park, Whiyoung Jung, Seungyul Han, S. K. Choi, Youngchul Sung
SJR Q1Neurocomputing
Computational Theory and MathematicsComputer Science
12
논문|인용수 1·2019
AMBER: Adaptive multi-batch experience replay for continuous action control
Seungyul Han, Youngchul Sung
Scholarworks@UNIST (Ulsan National Institute of Science and Technology)

In this paper, a new adaptive multi-batch experience replay scheme is proposed for proximal policy optimization (PPO) for continuous action control. On the contrary to original PPO, the proposed scheme uses the batch samples of past policies as well as the current policy for the update for the next policy, where the number of the used past batches is adaptively determined based on the oldness of the past batches measured by the average importance sampling (IS) weight. The new algorithm construct

Artificial IntelligenceComputer Science
13
논문|인용수 0·2024
Exclusively Penalized Q-learning for Offline Reinforcement Learning
Seungyul Han, Yonghyeon Jo, Jungmo Kim, Sanghyeon Lee, Junghyuk Yeom
Artificial IntelligenceComputer Science
15
논문|인용수 0·2026
Generalized Per-Agent Advantage Estimation for Multi-Agent Policy Optimization
Seongmin Kim, Giseung Park, Woojun Kim, Jiwon Jeon, Seungyul Han, Youngchul Sung
ArXiv.orgOA

In this paper, we propose a novel framework for multi-agent reinforcement learning that enhances sample efficiency and coordination through accurate per-agent advantage estimation. The core of our approach is Generalized Per-Agent Advantage Estimator (GPAE), which employs a per-agent value iteration operator to compute precise per-agent advantages. This operator enables stable off-policy learning by indirectly estimating values via action probabilities, eliminating the need for direct Q-function

Artificial IntelligenceComputer Science

대표 연구 분야

Artificial IntelligenceElectrical and Electronic EngineeringComputational Theory and MathematicsComputer Networks and Communications

한승열 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.