Korea Advanced Institute of Science and Technology · 工学
Professor Sangyeob Kim's research lab specializes in energy-efficient artificial intelligence hardware, focusing on ultra-low power neuromorphic computing and deep learning processors. The lab develops innovative architectures for spiking neural networks (SNNs), convolutional neural networks (CNNs), and transformer-based large language models (LLMs), emphasizing hardware-software co-design to minimize power consumption and memory access. Key research directions include on-chip learning, weight pruning, and memory-efficient inference through novel circuit techniques such as sign-extended bit gating and 1-bit comparators. The lab also investigates sensor-integrated systems for real-time signal processing, particularly in dynamic environments like sloshing fluid dynamics.
Figures are computed from collected data and may differ slightly.
This study considers a comparative study on pressure sensors for the measurement of sloshing impact pressure. For the comparative study, four pressure sensors are used: one piezoresistive sensor, one piezoelectric sensor, and two integrated circuit piezoelectric (ICP) sensors. For the comparative study, the sensors are installed on tank wall and ceiling of a rectangular tank with narrow breadth. Several types of comparative studies are carried out, including the sensitivity to temperature differ
Spiking-Neural-Networks (SNNs) have been studied for a long time, and recently have been shown to achieve the same accuracy as Convolutional-Neural-Networks (CNNs). By using CNN-to-SNN conversion, SNNs become a promising candidate for ultra-low power Al applications [1]. For example, compared to BNNs or XOR-nets, SNNs provide lower power consumption and higher accuracy [2]. This is because SNNs perform spike-based event-driven operation with high spike sparsity, unlike a CNN's frame-driven opera
A highly energy-efficient neuromorphic computing-in-memory (Neuro-CIM) processor is proposed for ultralow-power deep learning applications. Neuro-CIM can support spiking neural network (SNN) to eliminate the power and area overhead of previous CIM processor. The sign extended bits gating reduces the bitline (BL) voltage switching rate due to negative small-magnitude weights allowing 38% power reduction at 8-b weight condition and 25% at 4-b weight condition. In addition, Neuro-CIM replaces high-
An energy efficient Deep-Neural-Network (DNN) learning processor is proposed for on-chip learning and iterative weight pruning (WP). This work has three key features: 1) stochastic coarse-fine pruning reduced computation workload by 99.7% compared with previous WP algorithm while maintaining high weight sparsity, 2) adaptive input/output/weight skipping (AIOWS) achieved 30.1× higher throughput than previous DNN learning processor [1] for not only the inference but also learning, 3) weight memory
Recently, transformer-based large language models (LLMs), shown in Fig. 20.5.1, are widely used, and even on-device LLM systems with real-time responses are anticipated [1]. Many transformer processors [2–4] enhance energy efficiency by increasing hardware utilization and reducing power consumption, but their system power consumption and response time are still not suitable for mobile devices. Since LLMs, such as GPT-2, have many parameters (400-700M), their External Memory Access (EMA) consumes
A low power face recognition (FR) convolutional neural network (CNN) processor is proposed with high power efficiency to achieve always-on FR in mobile devices. Three key features enable a power-efficient FR CNN. First, tile-based clustering (THC) is proposed for reducing the computation overhead of hierarchical clustering. It generates an average of 37.2% duplicated input features in the entire network. Second, a low latency tile-based hierarchical clustering core is proposed. It supports an ap
In this article, we propose a complementary deep-neural-network (C-DNN) processor by combining convolutional neural network (CNN) and spiking neural network (SNN) to take advantage of them. The C-DNN processor can support both complementary inference and training with heterogeneous CNN and SNN core architecture. In addition, the C-DNN processor is the first DNN accelerator application-specific integrated circuit (ASIC) that can support CNN–SNN workload division by using their magnitude–energy tr
This paper considers scale effects on three-dimensional (3D) sloshing flows. A series of model tests were conducted for three differently scaled tanks. The model tanks considered in this study were 1:70, 1:50, and 1:30 scaled membrane type tanks based on a 138,000 m3 liquid natural gas carrier model. The tests were carried out for harmonic sway and roll motions for three different filling depths and with various excitation frequencies. The pressure measuring points in the tanks were the same, as
An energy-efficient neuromorphic computing-in-memory (CIM) processor is proposed with four key features: 1) Most significant bit (MSB) Word Skipping to reduce the BL activity; 2) Early Stopping to enable lower BL activity; 3) Mixed-mode firing for multi-macro aggregation; 4) Voltage Folding to extend the dynamic range. The proposed CIM achieves state-of-the-art energy efficiency of 62.1 TOPS/W (I=4b, W=8b) and 310.4 TOPS/W (I=4b, W=1b).
In this article, an energy-efficient spike domain deep-neural-network processor (SNPU) is proposed. Recently, many sparsity-aware processors were proposed to increase energy efficiency. However, because the activation of the real-world network was not high compared to the ideal condition, they were unable to completely employ integrated zero skipping logic. In addition, they employed weight sparsity by pruning, but their zero skipping logic was designed to perform best only under specific sparsi
This article proposes the TSUNAMI, which supports an energy-efficient deep-neural-network training. The TSUNAMI supports multi-modal iterative pruning to generate zeros in activation and weight. Tile-based dynamic activation pruning unit and weight memory shared pruning unit eliminate additional memory access. Coarse-zero skipping controller skips multiple unnecessary multiply-and-accumulation (MAC) operations at once, and fine-zero skipping controller skips randomly located unnecessary MAC oper
Recently, deep-neural-network (DNN) learning processors for edge devices have been proposed, but they cannot reduce the complexity of over-parameterized network during training. Also, they cannot support energy-efficient zero-skipping because previous methods cannot be performed perfectly in backpropagation and weight gradient update. In this letter, energy-efficient DNN learning processor PNPU is proposed with three key features: 1) stochastic coarse-fine level pruning; 2) adaptive input, outpu
This paper presents a numerical and experimental study of sloshing loads on liquefied natural gas (LNG) vessels. Conventional LNG carriers with membrane-type cargo systems have filling restrictions from 10% to 70% of tank height. The main reason for such restrictions is high sloshing loads around these filling depths. However, intermediate filling depths cannot be avoided for most LNG vessels except the LNG carrier. This study attempted to design a membrane-type LNG tank with a modified lower-ch
Recently, multiple ASICs [1]–[6] have been proposed to accelerate large language models (LLMs). However, the enormous number of LLM parameters leads to significant energy consumption due to external memory access (EMA). When normalizing the system energy required to process 1024 input tokens by the number of parameters, previous ASICs [1]–[5] required 79-222pJ/param for small models with 336-682M parameters, as shown in Fig. 23.9.1. Even an ASIC [6] designed to reduce EMA still consumes 26pJ/par
Open papers in the app to read, cite, and organize with AI.