Juhyoung Lee
KAIST 전산학과 · 컴퓨터과학
이 교수의 연구실은 에너지 효율적인 신경망 처리 아키텍처 설계에 초점을 맞추고 있으며, 특히 모바일 및 엣지 디바이스에 최적화된 초해상도 이미지 복원(SR)과 딥 강화학습(DRL) 프로세서의 고성능·저전력 구현을 핵심 연구 방향으로 삼고 있습니다. 컴퓨팅 인 메모리(CIM) 기반의 병렬 처리 아키텍처와 비트맵 압축, 선택적 캐싱, 스파arsity 기반 데이터 압축 기법을 통해 메모리 대역폭과 전력 소모를 극도로 줄이는 기술을 개발하고 있습니다. 특히 bfloat16 정밀도 기반의 고속 DNN 훈련 및 초해상도 복원 프로세서의 실현 가능성을 높이는 하드웨어-소프트웨어 통합 설계가 두드러집니다.
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
In this article, we propose an energy-efficient convolutional neural network (CNN) based super-resolution (SR) processor, super-resolution neural processing unit (SRNPU), for mobile applications. Traditionally, it is hard to realize real-time CNN-based SR on resource-limited platforms like mobile devices due to its massive amount of computation workload and communication bandwidth with external memory. The SRNPU can support the tile-based selective super-resolution (TSSR) which dynamically selec
An energy-efficient floating-point DNN training processor is proposed with heterogenous bfloat16 computing architecture using exponent computing-in-memory (CIM) and mantissa processing engine. Mantissa free exponent calculation enables pipelining of exponent and mantissa operation for heterogenous bfloat16 computing while reducing MAC power by 14.4 %. 6T SRAM exponent computing-in-memory with bitline charge reusing reduces memory access power by 46.4 %. The processor fabricated in 28 nm CMOS tec
A high-throughput CNN super resolution (SR) processor is proposed for memory efficient SR processing. It has three key features: 1) selective caching based layer fusion to minimize external memory access (EMA), 2) memory compaction scheme for smaller on-chip memory footprint, and 3) cyclic ring core architecture to increase the throughput with improved core utilization. As a result, the implemented processor achieves 60 frames-per-second throughput in generating full HD images.
This paper presents OmniDRL, a 4.18 TFLOPS and 29.3 TFLOPS/W DRL processor. A group-sparse training core and exponent mean delta encoding are proposed to enable weight and feature map compression for every iteration of DRL training. A sparse weight transposer enables on-chip transpose of compressed weight for reducing external memory access. The processor fabricated in 28 nm CMOS technology and occupies 3.6×3.6 mm <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org
The authors propose a heterogeneous floating-point (FP) computing architecture to maximize energy efficiency by separately optimizing exponent processing and mantissa processing. The proposed exponent-computing-in-memory architecture and mantissa-free exponent-computing algorithm reduce the power consumption of both memory and FP MAC while resolving previous FP computing-in-memory processors’ limitations. Also, a bfloat16 DNN training processor with proposed features and sparsity exploitation su
Deep reinforcement learning (DRL) has shown remarkable success in sequential decision-making problems but suffers from a long training time to obtain such good performance. Many parallel and distributed DRL training approaches have been proposed to solve this problem, but it is difficult to utilize them on resource-limited devices. In order to accelerate DRL in real-world edge devices, memory bandwidth bottlenecks due to large weight transactions have to be resolved. However, previous iterative
In this article, we present an energy-efficient deep reinforcement learning (DRL) processor, OmniDRL, for DRL training on edge devices. Recently, the need for DRL training is growing due to the DRL’s distinct characteristics that can be adapted to each user. However, a massive amount of external and internal memory access limits the implementation of DRL training on resource-constrained platforms. OmniDRL proposes four key features that can reduce external memory access by compressing as much da
Deep reinforcement learning (DRL) is widely used for autonomous systems including autonomous driving, robots, and drones. DRL training is essential for human-level control and adaptation to rapidly changing environments in mobile autonomous systems. However, acceleration of DRL training has three challenges: 1) large memory access, 2) various data patterns, 3) complex data dependency due to utilization of multiple DNNs. Two CMOS DRL accelerators have been proposed to support high speed, high ene
Recently, a lot of autonomous systems such as self-driving have been as close to human-level due to the rapid improvement of deep neural networks (DNN). However, they only performed well in the pre-trained environment and cannot achieve the same performance for the sudden environmental changes. To resolve this issue, self-adaptation through deep reinforcement learning (DRL) has been highlighted. However, it is hard to utilize DRL in the autonomous system due to the massive memory bandwidth requi
We define that the end effector is the device which interacts with the environment or contacts objects to execute tasks. Up to now, many researchers have developed anthropomorphic robotic hands as end effectors. We discuss the problems in the development of a human-scale and motor-driven anthropomorphic robot hand. In this paper, PRH's design concept, kinematic design and developments of the actuator, transmission, and sensing device are presented. By imitating the physiology of human hands, we
Global aviation is projected to grow in demand by an annual average of 4.1% between 2014 and 2034. It can be said that environmental impact from aviation will therefore be expected to increase on a similar scale. As regards civil aviation emissions, the sector contributes between 2~3% to International aviation GHG emissions. In the European Union(EU), aviation emissions account for about 3% of the EU's total green house gas emissions, of which a majority are said to come from international fligh
Abstract of Proposed FP CIM Processor (1) Heterogeneous FP Computing Arch. : Separate optimization of FP computing: Realize 2 cycles FP MAC w/ CIM (2) Exponent Computing-in-Memory: In-memory AND/NOR + BL charge reusing: Total memory power 46.4% 2) Mantissa Free Exponent Calculation: Removing redundant normalization: Total MAC power 14.4%
A real-time global matching optical flow estimation (OFE) processor is proposed for action recognition in mobile devices. The global OFE requires a large number of external memory accesses (EMAs) and matrix computations, thus it is incompatible on mobile devices with real-time constraints. For real-time OFE on mobile devices, this paper proposes two key features, both of which to reduce the required memory bandwidth and a number of computations: 1) Tile-based hierarchical OFE enables intermediat
본 연구는 선행연구에서 제안한 ‘디자인 프로세스 기반의 친환경 지속가능 디자인 가이드라인’을 사례 연구를 통해 평가하고, 분야별 활용 제안을 목적으로 한다. 선행 연구에서는 지속가능 디자인 요소들과 제품 개발 프로세스 단계별로 분석하고 이를 개발 프로세스에 해당 단계에 배치하여, 지속가능 디자인 가이드라인을 제안하였다. 본 연구에서는 개발된 지속가능 디자인 가이드라인을 평가하고, 디자인 분야별 활용 가능성을 제시하기 위해 디자인 사례를 선정하여 분석하였다. 분석한 결과 지속가능한 디자인 가이드라인에 누락된 요소를 발견할 수 있었다. 누락된 요소는 ‘자원 낭비 최소화’, ‘제작의 용이성’, ‘생산 자투리 소재 사용’, ‘사용의 편의성’, ‘사용자 생산 에너지 사용’, ‘운송의 안정성 강화’ 여섯 가지이다. 이를 추가하여, 최종 지속가능 디자인을 위한 가이드라인과 활용 사례를 제안하였다. 지속가능 디자인 가이드라인의 특징으로는 디자인 분야별 통합적 활용가능성과 디자인 실무자들의 지속가능