Skip to main content

Seungyul Han

Ulsan National Institute of Science and Technology · 情報科学

研究室紹介

Professor Seungyul Han's research lab specializes in advanced reinforcement learning methodologies with a focus on scalable and efficient algorithms for complex decision-making problems. The lab explores deep reinforcement learning, multi-agent reinforcement learning, and Lyapunov optimization to address challenges in exploration, policy learning, and system stability. Key research directions include sample-efficient learning, off-policy training with experience replay, and entropy regularization techniques for improved exploration in high-dimensional and partially observable environments. The lab also investigates practical applications in communication systems and dynamic resource management.

reinforcement learningmulti-agent RLoff-policy learningexploration efficiencyLyapunov optimization

Research Overview

Papers
40
Total Citations
116
Papers (5y)
30
Primary Field
情報科学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
30total
2022
2023
2024
2025
2026
Citations per year (5y)
12total
20222023202420252026

Selected Papers

15
1
Article|60 citations·2016
Filter Design for Generalized Frequency-Division Multiplexing
Seungyul Han, Youngchul Sung, Yong H. Lee
SJR Q1IEEE Transactions on Signal ProcessingOA

In this paper, optimal filter design for generalized frequency-division multiplexing (GFDM) is considered under two design criteria: rate maximization and out-of-band (OOB) emission minimization. First, the problem of GFDM filter optimization for rate maximization is formulated by expressing the transmission rate of GFDM as a function of GFDM filter coefficients. It is shown that Dirichlet filters are rate-optimal in additive white Gaussian noise channels with no carrier frequency offset (CFO) u

Electrical and Electronic EngineeringEngineering
2
Preprint|10 citations·2020
Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration
Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

In this paper, sample-aware policy entropy regularization is proposed to enhance the conventional policy entropy regularization for better exploration. Exploiting the sample distribution obtainable from the replay buffer, the proposed sample-aware entropy regularization maximizes the entropy of the weighted sum of the policy action distribution and the sample action distribution from the replay buffer for sample-efficient exploration. A practical algorithm named diversity actor-critic (DAC) is d

Artificial IntelligenceComputer Science
3
Preprint|9 citations·2020
A Reinforcement Learning Formulation of the Lyapunov Optimization: Application to Edge Computing Systems with Queue Stability
Sohee Bae, Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

In this paper, a deep reinforcement learning (DRL)-based approach to the Lyapunov optimization is considered to minimize the time-average penalty while maintaining queue stability. A proper construction of state and action spaces is provided to form a proper Markov decision process (MDP) for the Lyapunov optimization. A condition for the reward function of reinforcement learning (RL) for queue stability is derived. Based on the analysis and practical RL with reward discounting, a class of reward

Computer Networks and CommunicationsComputer Science
4
Article|7 citations·2024
FoX: Formation-Aware Exploration in Multi-Agent Reinforcement Learning
Yonghyeon Jo, Sunwoo Lee, Junghyuk Yeom, Seungyul Han
Proceedings of the AAAI Conference on Artificial IntelligenceOA

Recently, deep multi-agent reinforcement learning (MARL) has gained significant popularity due to its success in various cooperative multi-agent tasks. However, exploration still remains a challenging problem in MARL due to the partial observability of the agents and the exploration space that can grow exponentially as the number of agents increases. Firstly, in order to address the scalability issue of the exploration space, we define a formation-based equivalence relation on the exploration sp

Artificial IntelligenceComputer Science
5
Article|6 citations·2019
Dimension-Wise Importance Sampling Weight Clipping for Sample-Efficient Reinforcement Learning
Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

In importance sampling (IS)-based reinforcement learning algorithms such as Proximal Policy Optimization (PPO), IS weights are typically clipped to avoid large variance in learning. However, policy update from clipped statistics induces large bias in tasks with high action dimensions, and bias from clipping makes it difficult to reuse old samples with large IS weights. In this paper, we consider PPO, a representative on-policy algorithm, and propose its improvement by dimension-wise IS weight cl

Artificial IntelligenceComputer Science
6
Preprint|5 citations·2017
Multi-Batch Experience Replay for Fast Convergence of Continuous Action Control.
Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

Policy gradient methods for direct policy optimization are widely considered to obtain optimal policies in continuous Markov decision process (MDP) environments. However, policy gradient methods require exponentially many samples as the dimension of the action space increases. Thus, off-policy learning with experience replay is proposed to enable the agent to learn by using samples of other policies. Generally, large replay memories are preferred to minimize the sample correlation but large repl

Artificial IntelligenceComputer Science
7
Preprint|5 citations·2017
AMBER: Adaptive Multi-Batch Experience Replay for Continuous Action Control
Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

In this paper, a new adaptive multi-batch experience replay scheme is proposed for proximal policy optimization (PPO) for continuous action control. On the contrary to original PPO, the proposed scheme uses the batch samples of past policies as well as the current policy for the update for the next policy, where the number of the used past batches is adaptively determined based on the oldness of the past batches measured by the average importance sampling (IS) weight. The new algorithm construct

Artificial IntelligenceComputer Science
8
Preprint|5 citations·2021
A Max-Min Entropy Framework for Reinforcement Learning
Seungyul Han, Youngchul Sung
arXiv (Cornell University)OA

In this paper, we propose a max-min entropy framework for reinforcement learning (RL) to overcome the limitation of the soft actor-critic (SAC) algorithm implementing the maximum entropy RL in model-free sample-based learning. Whereas the maximum entropy RL guides learning for policies to reach states with high entropy in the future, the proposed max-min entropy framework aims to learn to visit states with low entropy and maximize the entropy of these low-entropy states to promote better explora

Artificial IntelligenceComputer Science
9
Preprint|3 citations·2022
Robust Imitation Learning against Variations in Environment Dynamics
Jong-Seong Chae, Seungyul Han, Whiyoung Jung, Myungsik Cho, Sungho Choi, Youngchul Sung
arXiv (Cornell University)OA

In this paper, we propose a robust imitation learning (IL) framework that improves the robustness of IL when environment dynamics are perturbed. The existing IL framework trained in a single environment can catastrophically fail with perturbations in environment dynamics because it does not capture the situation that underlying environment dynamics can be changed. Our framework effectively deals with environments with varying dynamics by imitating multiple experts in sampled environment dynamics

Artificial IntelligenceComputer Science
10
Preprint|3 citations·2020
Cross-Domain Imitation Learning with a Dual Structure
Sungho Choi, Seungyul Han, Woojun Kim, Youngchul Sung
arXiv (Cornell University)OA

In this paper, we consider cross-domain imitation learning (CDIL) in which an agent in a target domain learns a policy to perform well in the target domain by observing expert demonstrations in a source domain without accessing any reward function. In order to overcome the domain difference for imitation learning, we propose a dual-structured learning method. The proposed learning method extracts two feature vectors from each input observation such that one vector contains domain information and

Artificial IntelligenceComputer Science
11
Article|2 citations·2025
Adaptive multi-model fusion learning for sparse-reward reinforcement learning
Giseung Park, Whiyoung Jung, Seungyul Han, S. K. Choi, Youngchul Sung
SJR Q1Neurocomputing
Computational Theory and MathematicsComputer Science
12
Article|1 citations·2019
AMBER: Adaptive multi-batch experience replay for continuous action control
Seungyul Han, Youngchul Sung
Scholarworks@UNIST (Ulsan National Institute of Science and Technology)

In this paper, a new adaptive multi-batch experience replay scheme is proposed for proximal policy optimization (PPO) for continuous action control. On the contrary to original PPO, the proposed scheme uses the batch samples of past policies as well as the current policy for the update for the next policy, where the number of the used past batches is adaptively determined based on the oldness of the past batches measured by the average importance sampling (IS) weight. The new algorithm construct

Artificial IntelligenceComputer Science
13
Article|0 citations·2024
Exclusively Penalized Q-learning for Offline Reinforcement Learning
Seungyul Han, Yonghyeon Jo, Jungmo Kim, Sanghyeon Lee, Junghyuk Yeom
Artificial IntelligenceComputer Science
15
Article|0 citations·2026
Generalized Per-Agent Advantage Estimation for Multi-Agent Policy Optimization
Seongmin Kim, Giseung Park, Woojun Kim, Jiwon Jeon, Seungyul Han, Youngchul Sung
ArXiv.orgOA

In this paper, we propose a novel framework for multi-agent reinforcement learning that enhances sample efficiency and coordination through accurate per-agent advantage estimation. The core of our approach is Generalized Per-Agent Advantage Estimator (GPAE), which employs a per-agent value iteration operator to compute precise per-agent advantages. This operator enables stable off-policy learning by indirectly estimating values via action probabilities, eliminating the need for direct Q-function

Artificial IntelligenceComputer Science

Research Areas

Artificial IntelligenceElectrical and Electronic EngineeringComputational Theory and MathematicsComputer Networks and Communications

Seungyul Hanの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。