Jooyoung Kim
Korea Advanced Institute of Science and Technology · 工学
研究室紹介
Professor Jooyoung Kim's research lab specializes in computer architecture and hardware acceleration for emerging workloads, with a strong focus on optimizing datacenter-scale systems for AI and machine learning workloads. The lab develops custom, reconfigurable hardware fabrics using FPGAs to accelerate critical services such as web search ranking and natural language processing. Key research directions include memory-centric computing, low-power real-time perception systems, and domain-specific accelerators for vision and NLP workloads. The lab also explores innovative circuit-level designs, such as bitwise logic and energy-efficient comparators, to enhance performance and area efficiency in hardware accelerators.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15To advance datacenter capabilities beyond what commodity server designs can provide, the authors designed and built a composable, reconfigurable fabric to accelerate large-scale software services. Each instantiation of the fabric consists of a 6 x 8 2D torus of high-end field-programmable gate arrays (FPGAs) embedded into a half-rack of 48 servers. The authors deployed the reconfigurable fabric in a bed of 1,632 servers and FPGAs in a production datacenter and successfully used it to accelerate
A 201.4 GOPS real-time multi-object recognition processor is presented with a three-stage pipelined architecture. Visual perception based multi-object recognition algorithm is applied to give multiple attentions to multiple objects in the input image. For human-like multi-object perception, a neural perception engine is proposed with biologically inspired neural networks and fuzzy logic circuits. In the proposed hardware architecture, three recognition tasks (visual perception, descriptor genera
In these days, a technology that utilize of Bluetooth Low Energy (BLE) beacon, has been attracted attention to provide variety of convenience services. Especially, not limited to the service that can assist to people directly such as public safety, healthcare, proximity-based service, mobile payment, etc., a technology that can provide convenience indirectly such as asset tracking has also been proposed. Most of all, the indoor location awareness using BLE beacon is the essential technique that
Transformer is a deep learning language model widely used for natural language processing (NLP) services in datacenters. Among transformer models, Generative Pretrained Transformer (GPT) has achieved remarkable performance in text generation, or natural language generation (NLG), which needs the processing of a large input context in the summarization stage, followed by the generation stage that produces a single word at a time. The conventional platforms such as GPU are specialized for the para
Artificial intelligence (AI) and machine learning (ML) are revolutionizing many fields of study, such as visual recognition, natural language processing, autonomous vehicles, and prediction. Traditional von-Neumann computing architecture with separated processing elements and memory devices have been improving their computing performances rapidly with the scaling of process technology. However, in the era of AI and ML, data transfer between memory devices and processing elements becomes the bott
In this paper, we present a Bitwise Competition Logic (BCL) for the high performance and area efficient digital comparator. It compares two integer numbers using the location of the first 1 from the MSB, without arithmetic computations. The detail circuits to implement BCL, pre-encoder and selection logics are explained. The implemented BCL comparator shows 16%, 38% and 30% improved result in propagation delay, transistor count, and physical area compared to the other types of comparators. Measu
A 118.4 GB/s multi-casting network-on-chip (MC-NoC) is proposed as communication platform for a real-time object recognition processor. For application-specific NoC design, target traffic patterns are elaborately analyzed. Through topology exploration, we derive a hierarchical star and ring (HS-R) combined architecture for low latency and inter-processor communication. Multi-casting protocol and router are developed to accelerate one-to-many (1-to-N) data transactions. With these two main featur
We propose a novel location finding system exploiting the downlink reference signal of the Long Term Evolution (LTE) because of the absence of an efficient location finding system for the LTE. The proposed system is based on the correlation method and Chan's method for location sensing and positioning process, respectively. The conventional correlation method, however, is not matched to the LTE since the orthogonal gold sequence used for reference signal is assigned in frequency domain. Therefor
The visual attention mechanism, which is the way humans perform object recognition [1], was applied to the implementation of a high performance object recognition chip [2]. Even though the previous chip achieved 50% gain of computational cost [2], it could recognize only one object in a frame so that it is not suitable for advanced multi-object recognition applications such as video surveillance, intelligent robots, and autonomous vehicle navigation [3].
A proposed object recognition processor lightens its workload by estimating global region-of-interest features. A neuro-fuzzy controller performs intelligent ROI estimation by mimicking the human visual system, then manages the processor's overall pipeline stages using workload-aware task scheduling and applied database size control. The NFC performs workload-aware dynamic power management to reduce the proposed processor's power consumption.
A 66 fps 38 mW nearest neighbor matching processor for real-time object recognition has been fabricated in 0.13 mum CMOS technology. It consists of RISC processing core, pre-fetch DMA, and two independent sets of logic merged memories. Based on hierarchical vector quantization (H-VQ) algorithm, implemented processor achieves 22.5X cycle time reduction in matching process without any accuracy loss in VQ operation. As a result, 66 fps frame rate is obtained for QVGA (320times240 pixels) video imag
Abstract-Visual image processing random access memory (VIP-RAM) is proposed for a real-time multicore object recognition processor. It has two key features for the overall processor: 1) single cycle local maximum location search (LMLS) for fast key-point localization in object recognition, and 2) data consistency management (DCM) for producer-consumer data transactions among the processors. To achieve single cycle LMLS operation for a 3 x 3 window, the VIP-RAM adopts a hierarchical three-bank ar