Sangyeob Kim
KAIST 전기및전자공학부 · 공학
김상엽 교수의 연구실은 초저전력 인공지능 및 신경형 컴퓨팅 기반의 에너지 효율적 하드웨어 아키텍처 설계를 핵심으로 합니다. 압축된 신경망, 스파iking 신경망(SNN), 그리고 양자화 및 프루닝 기반의 딥러닝 프로세서 개발을 통해 모바일 및 임베디드 환경에서의 실시간 인공지능 처리를 구현하고자 합니다. 특히, 외부 메모리 접근을 최소화하고, 하드웨어 자원의 활용도를 극대화하는 기술적 혁신을 지속적으로 연구하고 있습니다.
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
This study considers a comparative study on pressure sensors for the measurement of sloshing impact pressure. For the comparative study, four pressure sensors are used: one piezoresistive sensor, one piezoelectric sensor, and two integrated circuit piezoelectric (ICP) sensors. For the comparative study, the sensors are installed on tank wall and ceiling of a rectangular tank with narrow breadth. Several types of comparative studies are carried out, including the sensitivity to temperature differ
Spiking-Neural-Networks (SNNs) have been studied for a long time, and recently have been shown to achieve the same accuracy as Convolutional-Neural-Networks (CNNs). By using CNN-to-SNN conversion, SNNs become a promising candidate for ultra-low power Al applications [1]. For example, compared to BNNs or XOR-nets, SNNs provide lower power consumption and higher accuracy [2]. This is because SNNs perform spike-based event-driven operation with high spike sparsity, unlike a CNN's frame-driven opera
A highly energy-efficient neuromorphic computing-in-memory (Neuro-CIM) processor is proposed for ultralow-power deep learning applications. Neuro-CIM can support spiking neural network (SNN) to eliminate the power and area overhead of previous CIM processor. The sign extended bits gating reduces the bitline (BL) voltage switching rate due to negative small-magnitude weights allowing 38% power reduction at 8-b weight condition and 25% at 4-b weight condition. In addition, Neuro-CIM replaces high-
An energy efficient Deep-Neural-Network (DNN) learning processor is proposed for on-chip learning and iterative weight pruning (WP). This work has three key features: 1) stochastic coarse-fine pruning reduced computation workload by 99.7% compared with previous WP algorithm while maintaining high weight sparsity, 2) adaptive input/output/weight skipping (AIOWS) achieved 30.1× higher throughput than previous DNN learning processor [1] for not only the inference but also learning, 3) weight memory
Recently, transformer-based large language models (LLMs), shown in Fig. 20.5.1, are widely used, and even on-device LLM systems with real-time responses are anticipated [1]. Many transformer processors [2–4] enhance energy efficiency by increasing hardware utilization and reducing power consumption, but their system power consumption and response time are still not suitable for mobile devices. Since LLMs, such as GPT-2, have many parameters (400-700M), their External Memory Access (EMA) consumes
A low power face recognition (FR) convolutional neural network (CNN) processor is proposed with high power efficiency to achieve always-on FR in mobile devices. Three key features enable a power-efficient FR CNN. First, tile-based clustering (THC) is proposed for reducing the computation overhead of hierarchical clustering. It generates an average of 37.2% duplicated input features in the entire network. Second, a low latency tile-based hierarchical clustering core is proposed. It supports an ap
In this article, we propose a complementary deep-neural-network (C-DNN) processor by combining convolutional neural network (CNN) and spiking neural network (SNN) to take advantage of them. The C-DNN processor can support both complementary inference and training with heterogeneous CNN and SNN core architecture. In addition, the C-DNN processor is the first DNN accelerator application-specific integrated circuit (ASIC) that can support CNN–SNN workload division by using their magnitude–energy tr
This paper considers scale effects on three-dimensional (3D) sloshing flows. A series of model tests were conducted for three differently scaled tanks. The model tanks considered in this study were 1:70, 1:50, and 1:30 scaled membrane type tanks based on a 138,000 m3 liquid natural gas carrier model. The tests were carried out for harmonic sway and roll motions for three different filling depths and with various excitation frequencies. The pressure measuring points in the tanks were the same, as
An energy-efficient neuromorphic computing-in-memory (CIM) processor is proposed with four key features: 1) Most significant bit (MSB) Word Skipping to reduce the BL activity; 2) Early Stopping to enable lower BL activity; 3) Mixed-mode firing for multi-macro aggregation; 4) Voltage Folding to extend the dynamic range. The proposed CIM achieves state-of-the-art energy efficiency of 62.1 TOPS/W (I=4b, W=8b) and 310.4 TOPS/W (I=4b, W=1b).
In this article, an energy-efficient spike domain deep-neural-network processor (SNPU) is proposed. Recently, many sparsity-aware processors were proposed to increase energy efficiency. However, because the activation of the real-world network was not high compared to the ideal condition, they were unable to completely employ integrated zero skipping logic. In addition, they employed weight sparsity by pruning, but their zero skipping logic was designed to perform best only under specific sparsi
This article proposes the TSUNAMI, which supports an energy-efficient deep-neural-network training. The TSUNAMI supports multi-modal iterative pruning to generate zeros in activation and weight. Tile-based dynamic activation pruning unit and weight memory shared pruning unit eliminate additional memory access. Coarse-zero skipping controller skips multiple unnecessary multiply-and-accumulation (MAC) operations at once, and fine-zero skipping controller skips randomly located unnecessary MAC oper
Recently, deep-neural-network (DNN) learning processors for edge devices have been proposed, but they cannot reduce the complexity of over-parameterized network during training. Also, they cannot support energy-efficient zero-skipping because previous methods cannot be performed perfectly in backpropagation and weight gradient update. In this letter, energy-efficient DNN learning processor PNPU is proposed with three key features: 1) stochastic coarse-fine level pruning; 2) adaptive input, outpu
This paper presents a numerical and experimental study of sloshing loads on liquefied natural gas (LNG) vessels. Conventional LNG carriers with membrane-type cargo systems have filling restrictions from 10% to 70% of tank height. The main reason for such restrictions is high sloshing loads around these filling depths. However, intermediate filling depths cannot be avoided for most LNG vessels except the LNG carrier. This study attempted to design a membrane-type LNG tank with a modified lower-ch
Recently, multiple ASICs [1]–[6] have been proposed to accelerate large language models (LLMs). However, the enormous number of LLM parameters leads to significant energy consumption due to external memory access (EMA). When normalizing the system energy required to process 1024 input tokens by the number of parameters, previous ASICs [1]–[5] required 79-222pJ/param for small models with 336-682M parameters, as shown in Fig. 23.9.1. Even an ASIC [6] designed to reduce EMA still consumes 26pJ/par