Korea Advanced Institute of Science and Technology · Computer Science
Professor Juhyoung Lee's research lab specializes in energy-efficient hardware accelerators for deep learning workloads, with a strong focus on optimizing computation and memory hierarchies for mobile and edge devices. The lab develops specialized processor architectures for convolutional neural networks (CNNs), deep reinforcement learning (DRL), and DNN training, emphasizing low-power, high-throughput, and memory-bandwidth-efficient designs. Key innovations include computing-in-memory (CIM) architectures, sparse computation techniques, and heterogeneous floating-point processing to enable real-time AI inference and training on resource-constrained platforms. The lab’s work bridges algorithmic efficiency with custom hardware design, targeting applications in mobile vision, robotics, and edge AI.
Figures are computed from collected data and may differ slightly.
In this article, we propose an energy-efficient convolutional neural network (CNN) based super-resolution (SR) processor, super-resolution neural processing unit (SRNPU), for mobile applications. Traditionally, it is hard to realize real-time CNN-based SR on resource-limited platforms like mobile devices due to its massive amount of computation workload and communication bandwidth with external memory. The SRNPU can support the tile-based selective super-resolution (TSSR) which dynamically selec
An energy-efficient floating-point DNN training processor is proposed with heterogenous bfloat16 computing architecture using exponent computing-in-memory (CIM) and mantissa processing engine. Mantissa free exponent calculation enables pipelining of exponent and mantissa operation for heterogenous bfloat16 computing while reducing MAC power by 14.4 %. 6T SRAM exponent computing-in-memory with bitline charge reusing reduces memory access power by 46.4 %. The processor fabricated in 28 nm CMOS tec
A high-throughput CNN super resolution (SR) processor is proposed for memory efficient SR processing. It has three key features: 1) selective caching based layer fusion to minimize external memory access (EMA), 2) memory compaction scheme for smaller on-chip memory footprint, and 3) cyclic ring core architecture to increase the throughput with improved core utilization. As a result, the implemented processor achieves 60 frames-per-second throughput in generating full HD images.
This paper presents OmniDRL, a 4.18 TFLOPS and 29.3 TFLOPS/W DRL processor. A group-sparse training core and exponent mean delta encoding are proposed to enable weight and feature map compression for every iteration of DRL training. A sparse weight transposer enables on-chip transpose of compressed weight for reducing external memory access. The processor fabricated in 28 nm CMOS technology and occupies 3.6×3.6 mm <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org
The authors propose a heterogeneous floating-point (FP) computing architecture to maximize energy efficiency by separately optimizing exponent processing and mantissa processing. The proposed exponent-computing-in-memory architecture and mantissa-free exponent-computing algorithm reduce the power consumption of both memory and FP MAC while resolving previous FP computing-in-memory processors’ limitations. Also, a bfloat16 DNN training processor with proposed features and sparsity exploitation su
Deep reinforcement learning (DRL) has shown remarkable success in sequential decision-making problems but suffers from a long training time to obtain such good performance. Many parallel and distributed DRL training approaches have been proposed to solve this problem, but it is difficult to utilize them on resource-limited devices. In order to accelerate DRL in real-world edge devices, memory bandwidth bottlenecks due to large weight transactions have to be resolved. However, previous iterative
In this article, we present an energy-efficient deep reinforcement learning (DRL) processor, OmniDRL, for DRL training on edge devices. Recently, the need for DRL training is growing due to the DRL’s distinct characteristics that can be adapted to each user. However, a massive amount of external and internal memory access limits the implementation of DRL training on resource-constrained platforms. OmniDRL proposes four key features that can reduce external memory access by compressing as much da
Deep reinforcement learning (DRL) is widely used for autonomous systems including autonomous driving, robots, and drones. DRL training is essential for human-level control and adaptation to rapidly changing environments in mobile autonomous systems. However, acceleration of DRL training has three challenges: 1) large memory access, 2) various data patterns, 3) complex data dependency due to utilization of multiple DNNs. Two CMOS DRL accelerators have been proposed to support high speed, high ene
Recently, a lot of autonomous systems such as self-driving have been as close to human-level due to the rapid improvement of deep neural networks (DNN). However, they only performed well in the pre-trained environment and cannot achieve the same performance for the sudden environmental changes. To resolve this issue, self-adaptation through deep reinforcement learning (DRL) has been highlighted. However, it is hard to utilize DRL in the autonomous system due to the massive memory bandwidth requi
We define that the end effector is the device which interacts with the environment or contacts objects to execute tasks. Up to now, many researchers have developed anthropomorphic robotic hands as end effectors. We discuss the problems in the development of a human-scale and motor-driven anthropomorphic robot hand. In this paper, PRH's design concept, kinematic design and developments of the actuator, transmission, and sensing device are presented. By imitating the physiology of human hands, we
Global aviation is projected to grow in demand by an annual average of 4.1% between 2014 and 2034. It can be said that environmental impact from aviation will therefore be expected to increase on a similar scale. As regards civil aviation emissions, the sector contributes between 2~3% to International aviation GHG emissions. In the European Union(EU), aviation emissions account for about 3% of the EU's total green house gas emissions, of which a majority are said to come from international fligh
Abstract of Proposed FP CIM Processor (1) Heterogeneous FP Computing Arch. : Separate optimization of FP computing: Realize 2 cycles FP MAC w/ CIM (2) Exponent Computing-in-Memory: In-memory AND/NOR + BL charge reusing: Total memory power 46.4% 2) Mantissa Free Exponent Calculation: Removing redundant normalization: Total MAC power 14.4%
A real-time global matching optical flow estimation (OFE) processor is proposed for action recognition in mobile devices. The global OFE requires a large number of external memory accesses (EMAs) and matrix computations, thus it is incompatible on mobile devices with real-time constraints. For real-time OFE on mobile devices, this paper proposes two key features, both of which to reduce the required memory bandwidth and a number of computations: 1) Tile-based hierarchical OFE enables intermediat
본 연구는 선행연구에서 제안한 ‘디자인 프로세스 기반의 친환경 지속가능 디자인 가이드라인’을 사례 연구를 통해 평가하고, 분야별 활용 제안을 목적으로 한다. 선행 연구에서는 지속가능 디자인 요소들과 제품 개발 프로세스 단계별로 분석하고 이를 개발 프로세스에 해당 단계에 배치하여, 지속가능 디자인 가이드라인을 제안하였다. 본 연구에서는 개발된 지속가능 디자인 가이드라인을 평가하고, 디자인 분야별 활용 가능성을 제시하기 위해 디자인 사례를 선정하여 분석하였다. 분석한 결과 지속가능한 디자인 가이드라인에 누락된 요소를 발견할 수 있었다. 누락된 요소는 ‘자원 낭비 최소화’, ‘제작의 용이성’, ‘생산 자투리 소재 사용’, ‘사용의 편의성’, ‘사용자 생산 에너지 사용’, ‘운송의 안정성 강화’ 여섯 가지이다. 이를 추가하여, 최종 지속가능 디자인을 위한 가이드라인과 활용 사례를 제안하였다. 지속가능 디자인 가이드라인의 특징으로는 디자인 분야별 통합적 활용가능성과 디자인 실무자들의 지속가능
Open papers in the app to read, cite, and organize with AI.