Skip to main content

심재웅 교수

Jae Woong Shim

서울대학교 · 컴퓨터과학

연구실 소개

심재웅 교수의 연구실은 고성능 컴퓨팅과 신경망 처리를 융합한 차세대 하드웨어 아키텍처 설계를 주요 연구 분야로 삼고 있습니다. 특히 FPGA 기반의 신경망 가속기, 이중화 신경망(BNN) 및 순환 신경망(GRU)의 하드웨어 최적화를 통해 에너지 효율성과 성능을 동시에 향상시키는 기술을 개발하고 있습니다. 또한, 메모리-기반 컴퓨팅(PIM)과 통합 메모리 아키텍처를 활용한 그래프 컴퓨팅 및 고성능 애플리케이션의 성능 향상 방안을 탐구하고 있습니다. 이는 실시간 데이터 분석과 에지 컴퓨팅 환경에서 필수적인 고속·저전력 하드웨어 솔루션을 확보하기 위함입니다.

FPGA 가속기신경망 최적화메모리 기반 컴퓨팅에너지 효율하이브리드 메모리 아키텍처

연구 현황

논문 수
59
총 인용 수
2,536
최근 5년 논문
24
주요 분야
컴퓨터과학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
24총합
2022
2023
2024
2025
2026
5개년 연도별 피인용 수
213총합
20222023202420252026

주요 논문

15
1
논문|인용수 448·2017
Can FPGAs Beat GPUs in Accelerating Next-Generation Deep Neural Networks?
Eriko Nurvitadhi, Ganesh Venkatesh, Jaewoong Sim, Debbie Marr, Randy Huang, Jason Ong Gee Hock, Yeong Tat Liew, Srivatsan Krishnan, Duncan J. M. Moss, Suchit Subhaschandra, Guy Boudoukh
FWCI 51.7

Current-generation Deep Neural Networks (DNNs), such as AlexNet and VGG, rely heavily on dense floating-point matrix multiplication (GEMM), which maps well to GPUs (regular parallelism, high TFLOP/s). Because of this, GPUs are widely used for accelerating DNNs. Current FPGAs offer superior energy efficiency (Ops/Watt), but they do not offer the performance of today's GPUs on DNNs. In this paper, we look at upcoming FPGA technology advances, the rapid pace of innovation in DNN algorithms, and con

Hardware and ArchitectureComputer Science
2
논문|인용수 337·2016
Accelerating Binarized Neural Networks: Comparison of FPGA, CPU, GPU, and ASIC
Eriko Nurvitadhi, David Sheffield, Jaewoong Sim, Asit Mishra, Ganesh Venkatesh, Debbie Marr
FWCI 13.9

Deep neural networks (DNNs) are widely used in data analytics, since they deliver state-of-the-art accuracies. Binarized neural networks (BNNs) are recently proposed optimized variant of DNNs. BNNs constraint network weight and/or neuron value to either +1 or −1, which is representable in 1 bit. This leads to dramatic algorithm efficiency improvement, due to reduction in the memory and computational demands. This paper evaluates the opportunity to further improve the execution efficiency of BNNs

Computer Vision and Pattern RecognitionComputer Science
3
논문|인용수 270·2017
GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks
Lifeng Nai, Ramyad Hadidi, Jaewoong Sim, Hyojong Kim, Pranith Kumar, Hyesoon Kim
FWCI 25.5

With the emergence of data science, graph computing has become increasingly important these days. Unfortunately, graph computing typically suffers from poor performance when mapped to modern computing systems because of the overhead of executing atomic operations and inefficient utilization of the memory subsystem. Meanwhile, emerging technologies, such as Hybrid Memory Cube (HMC), enable the processing-in-memory (PIM) functionality with offloading operations at an instruction level. Instruction

Hardware and ArchitectureComputer Science
4
논문|인용수 191·2012
A performance analysis framework for identifying potential benefits in GPGPU applications
Jaewoong Sim, Aniruddha Dasgupta, Hyesoon Kim, Richard Vuduc
FWCI 22.3

Tuning code for GPGPU and other emerging many-core platforms is a challenge because few models or tools can precisely pinpoint the root cause of performance bottlenecks. In this paper, we present a performance analysis framework that can help shed light on such bottlenecks for GPGPU applications. Although a handful of GPGPU profiling tools exist, most of the traditional tools, unfortunately, simply provide programmers with a variety of measurements and metrics obtained by running applications, a

Hardware and ArchitectureComputer Science
5
논문|인용수 173·2016
Accelerating recurrent neural networks in analytics servers: Comparison of FPGA, CPU, GPU, and ASIC
Eriko Nurvitadhi, Jaewoong Sim, David Sheffield, Asit Mishra, Srivatsan Krishnan, Debbie Marr
FWCI 20.3

Recurrent neural networks (RNNs) provide state-of-the-art accuracy for performing analytics on datasets with sequence (e.g., language model). This paper studied a state-of-the-art RNN variant, Gated Recurrent Unit (GRU). We first proposed memoization optimization to avoid 3 out of the 6 dense matrix vector multiplications (SGEMVs) that are the majority of the computation in GRU. Then, we study the opportunities to accelerate the remaining SGEMVs using FPGAs, in comparison to 14-nm ASIC, GPU, and

Hardware and ArchitectureComputer Science
6
논문|인용수 109·2014
Transparent Hardware Management of Stacked DRAM as Part of Memory
Jaewoong Sim, Alaa R. Alameldeen, Zeshan Chishti, Chris Wilkerson, Hyesoon Kim
FWCI 11.4

Recent technology advancements allow for the integration of large memory structures on-die or as a die-stacked DRAM. Such structures provide higher bandwidth and faster access time than off-chip memory. Prior work has investigated using the large integrated memory as a cache, or using it as part of a heterogeneous memory system under management of the OS. Using this memory as a cache would waste a large fraction of total memory space, especially for the systems where stacked memory could be as l

Hardware and ArchitectureComputer Science
7
논문|인용수 89·2012
A Mostly-Clean DRAM Cache for Effective Hit Speculation and Self-Balancing Dispatch
Jaewoong Sim, Gabriel H. Loh, Hyesoon Kim, Mike O’Connor, Mithuna Thottethodi
FWCI 7.7

Die-stacking technology allows conventional DRAM to be integrated with processors. While numerous opportunities to make use of such stacked DRAM exist, one promising way is to use it as a large cache. Although previous studies show that DRAM caches can deliver performance benefits, there remain inefficiencies as well as significant hardware costs for auxiliary structures. This paper presents two innovations that exploit the bursty nature of memory requests to streamline the DRAM cache. The first

Hardware and ArchitectureComputer Science
8
논문|인용수 82·2018
A Customizable Matrix Multiplication Framework for the Intel HARPv2 Xeon+FPGA Platform
Duncan J. M. Moss, Srivatsan Krishnan, Eriko Nurvitadhi, Piotr Ratuszniak, Christopher N. Johnson, Jaewoong Sim, Asit Mishra, Debbie Marr, Suchit Subhaschandra, Philip H. W. Leong
FWCI 15.3

General Matrix to Matrix multiplication (GEMM) is the cornerstone for a wide gamut of applications in high performance computing (HPC), scientific computing (SC) and more recently, deep learning. In this work, we present a customizable matrix multiplication framework for the Intel HARPv2 CPU+FPGA platform that includes support for both traditional single precision floating point and reduced precision workloads. Our framework supports arbitrary size GEMMs and consists of two parts: (1) a simple a

Hardware and ArchitectureComputer Science
9
논문|인용수 79·2020
Batch-Aware Unified Memory Management in GPUs for Irregular Workloads
Hyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi, Hyesoon Kim
FWCI 9.9

While unified virtual memory and demand paging in modern GPUs provide convenient abstractions to programmers for working with large-scale applications, they come at a significant performance cost. We provide the first comprehensive analysis of major inefficiencies that arise in page fault handling mechanisms employed in modern GPUs. To amortize the high costs in fault handling, the GPU runtime processes a large number of GPU page faults together. We observe that this batched processing of page f

Hardware and ArchitectureComputer Science
10
논문|인용수 69·2020
Active Learning of Convolutional Neural Network for Cost-Effective Wafer Map Pattern Classification
Jaewoong Shim, Seokho Kang, Sungzoon Cho
SJR Q2FWCI 7.4IEEE Transactions on Semiconductor Manufacturing

Wafer maps provide important information for engineers for detecting root causes of failure in a semiconductor manufacturing process. Thus, there has been active research into the automation of wafer map pattern classification. With recent advances in deep learning, a convolutional neural network (CNN) has yielded state-of-the-art performance in wafer map pattern classification. Because a large amount of labeled training data is required, experienced engineers need to annotate large quantities o

Industrial and Manufacturing EngineeringEngineering
11
논문|인용수 44·2013
Resilient die-stacked DRAM caches
Jaewoong Sim, Gabriel H. Loh, Vilas Sridharan, Mike O’Connor
FWCI 5.8

Die-stacked DRAM can provide large amounts of in-package, high-bandwidth cache storage. For server and high-performance computing markets, however, such DRAM caches must also provide sufficient support for reliability and fault tolerance. While conventional off-chip memory provides ECC support by adding one or more extra chips, this may not be practical in a 3D stack. In this paper, we present a DRAM cache organization that uses error-correcting codes (ECCs), strong checksums (CRCs), and dirty d

Electrical and Electronic EngineeringEngineering
12
논문|인용수 31·2012
FLEXclusion
Jaewoong Sim, Jaekyu Lee, Moinuddin K. Qureshi, Hyesoon Kim
FWCI 3.3ACM SIGARCH Computer Architecture News

Exclusive last-level caches (LLCs) reduce memory accesses by effectively utilizing cache capacity. However, they require excessive on-chip bandwidth to support frequent insertions of cache lines on eviction from upper-level caches. Non-inclusive caches, on the other hand, have the advantage of using the on-chip bandwidth more effectively but suffer from a higher miss rate. Traditionally, the decision to use the cache as exclusive or non-inclusive is made at design time. However, the best option

Hardware and ArchitectureComputer Science
13
논문|인용수 22·2021
Active cluster annotation for wafer map pattern classification in semiconductor manufacturing
Jaewoong Shim, Seokho Kang, Sungzoon Cho
SJR Q1FWCI 2.4Expert Systems with Applications
Industrial and Manufacturing EngineeringEngineering
14
논문|인용수 22·2023
Learning from single-defect wafer maps to classify mixed-defect wafer maps
Jaewoong Shim, Seokho Kang
SJR Q1FWCI 4.5Expert Systems with Applications
Industrial and Manufacturing EngineeringEngineering
15
논문|인용수 18·2021
Adaptive fault detection framework for recipe transition in semiconductor manufacturing
Jaewoong Shim, Sungzoon Cho, Euiseok Kum, Suho Jeong
SJR Q1FWCI 2.1Computers & Industrial Engineering
Industrial and Manufacturing EngineeringEngineering

대표 연구 분야

Hardware and ArchitectureIndustrial and Manufacturing EngineeringElectrical and Electronic EngineeringArtificial IntelligenceComputer Graphics and Computer-Aided DesignComputer Vision and Pattern Recognition

심재웅 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.