Skip to main content

박종세 교수

JongSeok Park

KAIST 전산학부 · 컴퓨터과학

연구실 소개

박종세 교수의 연구실은 딥러닝 및 머신러닝 워크로드의 고성능·고효율 처리를 위해 하드웨어 기반 가속 기술을 핵심으로 연구합니다. 특히 FPGA 기반의 자동 가속기 생성 프레임워크, 정밀도 저하에 강한 신경망 아키텍처 설계, 그리고 네트워크 인터페이스 카드 내장 압축 가속기 등을 통해 통신 비용과 에너지 소비를 극복하고자 합니다. 또한, 유사한 계산 특성을 가진 코드를 위한 아날로그 하드웨어 기반의 근사 계산 기법 등, 전통적인 von Neumann 아키텍처를 넘어서는 혁신적 접근도 함께 탐구하고 있습니다.

FPGA 가속기신경망 정밀도 최소화네트워크 내 압축 가속근사 계산하이브리드 아키텍처

연구 현황

논문 수
77
총 인용 수
2,275
최근 5년 논문
46
주요 분야
컴퓨터과학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
46총합
2022
2023
2024
2025
2026
5개년 연도별 피인용 수
254총합
20222023202420252026

주요 논문

15
1
논문|인용수 527·2018
Bit Fusion: Bit-Level Dynamically Composable Architecture for Accelerating Deep Neural Network
Hardik Sharma, Jongse Park, Naveen Suda, Liangzhen Lai, Benson Chau, Joon Kyung Kim, Vikas Chandra, Hadi Esmaeilzadeh

Hardware acceleration of Deep Neural Networks (DNNs) aims to tame their enormous compute intensity. Fully realizing the potential of acceleration in this domain requires understanding and leveraging algorithmic properties of DNNs. This paper builds upon the algorithmic insight that bitwidth of operations in DNNs can be reduced without compromising their classification accuracy. However, to prevent loss of accuracy, the bitwidth varies significantly across DNNs and it may even be adjusted for eac

Computer Vision and Pattern RecognitionComputer Science
2
논문|인용수 359·2016
From high-level deep neural models to FPGAs
Hardik Sharma, Jongse Park, Divya Mahajan, Emmanuel Amaro, Joon Kyung Kim, Chenkai Shao, Asit Mishra, Hadi Esmaeilzadeh

Deep Neural Networks (DNNs) are compute-intensive learning models with growing applicability in a wide range of domains. FPGAs are an attractive choice for DNNs since they offer a programmable substrate for acceleration and are becoming available across different market segments. However, obtaining both performance and energy efficiency with FPGAs is a laborious task even for expert hardware designers. Furthermore, the large memory footprint of DNNs, coupled with the FPGAs' limited on-chip stora

Computer Vision and Pattern RecognitionComputer Science
3
논문|인용수 160·2016
TABLA: A unified template-based framework for accelerating statistical machine learning
Divya Mahajan, Jongse Park, Emmanuel Amaro, Hardik Sharma, Amir Yazdanbakhsh, Joon Kyung Kim, Hadi Esmaeilzadeh
OA

A growing number of commercial and enterprise systems increasingly rely on compute-intensive Machine Learning (ML) algorithms. While the demand for these compute-intensive applications is growing, the performance benefits from general-purpose platforms are diminishing. Field Programmable Gate Arrays (FPGAs) provide a promising path forward to accommodate the needs of machine learning algorithms and represent an intermediate point between the efficiency of ASICs and the programmability of general

Computer Vision and Pattern RecognitionComputer Science
4
논문|인용수 148·2014
General-purpose code acceleration with limited-precision analog computation
Renée St. Amant, Amir Yazdanbakhsh, Jongse Park, Bradley Thwaites, Hadi Esmaeilzadeh, Arjang Hassibi, Luís Ceze, Doug Burger
ACM SIGARCH Computer Architecture News

As improvements in per-transistor speed and energy efficiency diminish, radical departures from conventional approaches are becoming critical to improving the performance and energy efficiency of general-purpose processors. We propose a solution--from circuit to compiler-that enables general-purpose use of limited-precision, analog hardwareto accelerate "approximable" code---code that can tolerate imprecise execution. We utilize an algorithmic transformation that automatically converts approxima

Hardware and ArchitectureComputer Science
5
논문|인용수 110·2015
Neural acceleration for GPU throughput processors
Amir Yazdanbakhsh, Jongse Park, Hardik Sharma, Pejman Lotfi-Kamran, Hadi Esmaeilzadeh
OA

Graphics Processing Units (GPUs) can accelerate diverse classes of applications, such as recognition, gaming, data analytics, weather prediction, and multimedia. Many of these applications are amenable to approximate execution. This application characteristic provides an opportunity to improve GPU performance and efficiency. Among approximation techniques, neural accelerators have been shown to provide significant performance and efficiency gains when augmenting CPU processors. However, the inte

Hardware and ArchitectureComputer Science
6
논문|인용수 90·2018
A Network-Centric Hardware/Algorithm Co-Design to Accelerate Distributed Training of Deep Neural Networks
Youjie Li, Jongse Park, Mohammad Alian, Yifan Yuan, Zheng Qu, Peitian Pan, Ren Wang, Alexander G. Schwing, Hadi Esmaeilzadeh, Nam Sung Kim

Training real-world Deep Neural Networks (DNNs) can take an eon (i.e., weeks or months) without leveraging distributed systems. Even distributed training takes inordinate time, of which a large fraction is spent in communicating weights and gradients over the network. State-of-the-art distributed training algorithms use a hierarchy of worker-aggregator nodes. The aggregators repeatedly receive gradient updates from their allocated group of the workers, and send back the updated weights. This pap

Computer Vision and Pattern RecognitionComputer Science
7
preprint|인용수 82·2024
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
Guseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi, S. K. Lee, Hyungkyu Ham, Gwangsun Kim, Divya Mahajan, Jongse Park
OA

Modern transformer-based Large Language Models (LLMs) are constructed with a series of decoder blocks. Each block comprises three key components: (1) QKV generation, (2) multi-head attention, and (3) feed-forward networks. In batched processing, QKV generation and feed-forward networks involve compute-intensive matrix-matrix multiplications (GEMM), while multi-head attention requires bandwidth-heavy matrix-vector multiplications (GEMV). Machine learning accelerators like TPUs or NPUs are profici

Electrical and Electronic EngineeringEngineering
8
논문|인용수 74·2015
FlexJava: language support for safe and modular approximate programming
Jongse Park, Hadi Esmaeilzadeh, Xin Zhang, Mayur Naik, William R. Harris

Energy efficiency is a primary constraint in modern systems. Approximate computing is a promising approach that trades quality of result for gains in efficiency and performance. State- of-the-art approximate programming models require extensive manual annotations on program data and operations to guarantee safe execution of approximate programs. The need for extensive manual annotations hinders the practical use of approximation techniques. This paper describes FlexJava, a small set of language

Hardware and ArchitectureComputer Science
9
논문|인용수 65·2014
General-purpose code acceleration with limited-precision analog computation
Renée St. Amant, Amir Yazdanbakhsh, Jongse Park, Bradley Thwaites, Hadi Esmaeilzadeh, Arjang Hassibi, Luís Ceze, Doug Burger

As improvements in per-transistor speed and energy efficiency diminish, radical departures from conventional approaches are becoming critical to improving the performance and energy efficiency of general-purpose processors. We propose a solution—from circuit to compiler—that enables general-purpose use of limited-precision, analog hardware to accelerate “approximable” code—code that can tolerate imprecise execution. We utilize an algorithmic transformation that automatically converts approximabl

Electrical and Electronic EngineeringEngineering
10
논문|인용수 64·2012
Locality-aware dynamic VM reconfiguration on MapReduce clouds
Jongse Park, Daewoo Lee, Bo-Kyeong Kim, Jaehyuk Huh, Seungryoul Maeng

Cloud computing based on system virtualization, has been expanding its services to distributed data-intensive platforms such as MapReduce and Hadoop. Such a distributed platform on clouds runs in a virtual cluster consisting of a number of virtual machines. In the virtual cluster, demands on computing resources for each node may fluctuate, due to data locality and task behavior. However, current cloud services use a static cluster configuration, fixing or manually adjusting the computing capabil

Information SystemsComputer Science
11
논문|인용수 49·2021
SLO-Aware Inference Scheduler for Heterogeneous Processors in Edge Platforms
Wonik Seo, Sang-Hoon Cha, Yeonjae Kim, Jaehyuk Huh, Jongse Park
SJR Q2ACM Transactions on Architecture and Code OptimizationOA

With the proliferation of applications with machine learning (ML), the importance of edge platforms has been growing to process streaming sensor, data locally without resorting to remote servers. Such edge platforms are commonly equipped with heterogeneous computing processors such as GPU, DSP, and other accelerators, but their computational and energy budget are severely constrained compared to the data center servers. However, as an edge platform must perform the processing of multiple machine

Computer Networks and CommunicationsComputer Science
12
논문|인용수 44·2015
Axilog: language support for approximate hardware design
Amir Yazdanbakhsh, Divya Mahajan, Bradley Thwaites, Jongse Park, Anandhavel Nagendrakumar, Sindhuja Sethuraman, Kartik Ramkrishnan, Nishanthi Ravindran, Rudra Jariwala, Abbas Rahimi, Hadi Esmaeilzadeh, Kia Bazargan
Design, Automation, and Test in Europe

Relaxing the traditional abstraction of near-perfect accuracy in hardware design can lead to significant gains in energy efficiency, area, and performance. To exploit this opportunity, there is a need for design abstractions that can systematically incorporate approximation in hardware design. We introduce Axilog, a set of language annotations, that provides the necessary syntax and semantics for approximate hardware design and reuse in Verilog. Axilog enables the designer to relax the accuracy

Electrical and Electronic EngineeringEngineering
13
논문|인용수 42·2015
Axilog: Language Support for Approximate Hardware Design
Amir Yazdanbakhsh, Divya Mahajan, Bradley Thwaites, Jongse Park, Anandhavel Nagendrakumar, Sindhuja Sethuraman, Kartik Ramkrishnan, Nishanthi Ravindran, Rudra Jariwala, Abbas Rahimi, Hadi Esmaeilzadeh, Kia Bazargan
Design, Automation & Test in Europe Conference & Exhibition (DATE), 2015

Relaxing the traditional abstraction of “near-perfect” accuracy in hardware design can lead to significant gains in energy efficiency, area, and performance. To exploit this opportunity, there is a need for design abstractions that can systematically incorporate approximation in hardware design. We introduce Axilog, a set of language annotations, that provides the necessary syntax and semantics for approximate hardware design and reuse in Verilog. Axilog enables the designer to relax the accurac

Hardware and ArchitectureComputer Science
14
논문|인용수 41·2014
Rollback-free value prediction with approximate loads
Bradley Thwaites, Gennady Pekhimenko, Hadi Esmaeilzadeh, Amir Yazdanbakhsh, Onur Mutlu, Jongse Park, Girish Mururu, Todd C. Mowry

This paper demonstrates how to utilize the inherent error resilience of a wide range of applications to mitigate the memory wall -- the discrepancy between core and memory speed. We define a new microarchitecturally-triggered approximation technique called rollback-free value prediction. This technique predicts the value of safe-to-approximate loads when they miss in the cache without tracking mispredictions or requiring costly recovery from misspeculations. This technique mitigates the memory w

Hardware and ArchitectureComputer Science
15
논문|인용수 41·2017
Scale-out acceleration for machine learning
Jongse Park, Hardik Sharma, Divya Mahajan, Joon Kyung Kim, Preston Olds, Hadi Esmaeilzadeh

The growing scale and complexity of Machine Learning (ML) algorithms has resulted in prevalent use of distributed general-purpose systems. In a rather disjoint effort, the community is focusing mostly on high performance single-node accelerators for learning. This work bridges these two paradigms and offers CoSMIC, a full computing stack constituting language, compiler, system software, template architecture, and circuit generators, that enable programmable acceleration of learning at scale. CoS

Hardware and ArchitectureComputer Science

대표 연구 분야

Hardware and ArchitectureComputer Vision and Pattern RecognitionElectrical and Electronic EngineeringArtificial IntelligenceComputer Networks and CommunicationsInformation Systems

박종세 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.