Skip to main content

JongSeok Park

Korea Advanced Institute of Science and Technology · 情報科学

研究室紹介

Professor JongSeok Park's research lab specializes in hardware-software co-design for energy-efficient and high-performance computing, with a strong focus on accelerating machine learning workloads—particularly Deep Neural Networks—through customized accelerators. The lab explores innovative solutions in FPGA-based acceleration, approximate computing using analog and low-precision computation, and in-network data compression to reduce communication bottlenecks in distributed training. Their work bridges algorithmic insights with low-level system design to enable efficient, scalable, and programmable acceleration for emerging AI workloads.

DNN accelerationFPGA accelerationapproximate computingin-network compressionhardware-software co-design

Research Overview

Papers
77
Total Citations
2,275
Papers (5y)
46
Primary Field
情報科学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
46total
2022
2023
2024
2025
2026
Citations per year (5y)
254total
20222023202420252026

Selected Papers

15
1
Article|527 citations·2018
Bit Fusion: Bit-Level Dynamically Composable Architecture for Accelerating Deep Neural Network
Hardik Sharma, Jongse Park, Naveen Suda, Liangzhen Lai, Benson Chau, Joon Kyung Kim, Vikas Chandra, Hadi Esmaeilzadeh

Hardware acceleration of Deep Neural Networks (DNNs) aims to tame their enormous compute intensity. Fully realizing the potential of acceleration in this domain requires understanding and leveraging algorithmic properties of DNNs. This paper builds upon the algorithmic insight that bitwidth of operations in DNNs can be reduced without compromising their classification accuracy. However, to prevent loss of accuracy, the bitwidth varies significantly across DNNs and it may even be adjusted for eac

Computer Vision and Pattern RecognitionComputer Science
2
Article|359 citations·2016
From high-level deep neural models to FPGAs
Hardik Sharma, Jongse Park, Divya Mahajan, Emmanuel Amaro, Joon Kyung Kim, Chenkai Shao, Asit Mishra, Hadi Esmaeilzadeh

Deep Neural Networks (DNNs) are compute-intensive learning models with growing applicability in a wide range of domains. FPGAs are an attractive choice for DNNs since they offer a programmable substrate for acceleration and are becoming available across different market segments. However, obtaining both performance and energy efficiency with FPGAs is a laborious task even for expert hardware designers. Furthermore, the large memory footprint of DNNs, coupled with the FPGAs' limited on-chip stora

Computer Vision and Pattern RecognitionComputer Science
3
Article|160 citations·2016
TABLA: A unified template-based framework for accelerating statistical machine learning
Divya Mahajan, Jongse Park, Emmanuel Amaro, Hardik Sharma, Amir Yazdanbakhsh, Joon Kyung Kim, Hadi Esmaeilzadeh
OA

A growing number of commercial and enterprise systems increasingly rely on compute-intensive Machine Learning (ML) algorithms. While the demand for these compute-intensive applications is growing, the performance benefits from general-purpose platforms are diminishing. Field Programmable Gate Arrays (FPGAs) provide a promising path forward to accommodate the needs of machine learning algorithms and represent an intermediate point between the efficiency of ASICs and the programmability of general

Computer Vision and Pattern RecognitionComputer Science
4
Article|148 citations·2014
General-purpose code acceleration with limited-precision analog computation
Renée St. Amant, Amir Yazdanbakhsh, Jongse Park, Bradley Thwaites, Hadi Esmaeilzadeh, Arjang Hassibi, Luís Ceze, Doug Burger
ACM SIGARCH Computer Architecture News

As improvements in per-transistor speed and energy efficiency diminish, radical departures from conventional approaches are becoming critical to improving the performance and energy efficiency of general-purpose processors. We propose a solution--from circuit to compiler-that enables general-purpose use of limited-precision, analog hardwareto accelerate "approximable" code---code that can tolerate imprecise execution. We utilize an algorithmic transformation that automatically converts approxima

Hardware and ArchitectureComputer Science
5
Article|110 citations·2015
Neural acceleration for GPU throughput processors
Amir Yazdanbakhsh, Jongse Park, Hardik Sharma, Pejman Lotfi-Kamran, Hadi Esmaeilzadeh
OA

Graphics Processing Units (GPUs) can accelerate diverse classes of applications, such as recognition, gaming, data analytics, weather prediction, and multimedia. Many of these applications are amenable to approximate execution. This application characteristic provides an opportunity to improve GPU performance and efficiency. Among approximation techniques, neural accelerators have been shown to provide significant performance and efficiency gains when augmenting CPU processors. However, the inte

Hardware and ArchitectureComputer Science
6
Article|90 citations·2018
A Network-Centric Hardware/Algorithm Co-Design to Accelerate Distributed Training of Deep Neural Networks
Youjie Li, Jongse Park, Mohammad Alian, Yifan Yuan, Zheng Qu, Peitian Pan, Ren Wang, Alexander G. Schwing, Hadi Esmaeilzadeh, Nam Sung Kim

Training real-world Deep Neural Networks (DNNs) can take an eon (i.e., weeks or months) without leveraging distributed systems. Even distributed training takes inordinate time, of which a large fraction is spent in communicating weights and gradients over the network. State-of-the-art distributed training algorithms use a hierarchy of worker-aggregator nodes. The aggregators repeatedly receive gradient updates from their allocated group of the workers, and send back the updated weights. This pap

Computer Vision and Pattern RecognitionComputer Science
7
Preprint|82 citations·2024
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
Guseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi, S. K. Lee, Hyungkyu Ham, Gwangsun Kim, Divya Mahajan, Jongse Park
OA

Modern transformer-based Large Language Models (LLMs) are constructed with a series of decoder blocks. Each block comprises three key components: (1) QKV generation, (2) multi-head attention, and (3) feed-forward networks. In batched processing, QKV generation and feed-forward networks involve compute-intensive matrix-matrix multiplications (GEMM), while multi-head attention requires bandwidth-heavy matrix-vector multiplications (GEMV). Machine learning accelerators like TPUs or NPUs are profici

Electrical and Electronic EngineeringEngineering
8
Article|74 citations·2015
FlexJava: language support for safe and modular approximate programming
Jongse Park, Hadi Esmaeilzadeh, Xin Zhang, Mayur Naik, William R. Harris

Energy efficiency is a primary constraint in modern systems. Approximate computing is a promising approach that trades quality of result for gains in efficiency and performance. State- of-the-art approximate programming models require extensive manual annotations on program data and operations to guarantee safe execution of approximate programs. The need for extensive manual annotations hinders the practical use of approximation techniques. This paper describes FlexJava, a small set of language

Hardware and ArchitectureComputer Science
9
Article|65 citations·2014
General-purpose code acceleration with limited-precision analog computation
Renée St. Amant, Amir Yazdanbakhsh, Jongse Park, Bradley Thwaites, Hadi Esmaeilzadeh, Arjang Hassibi, Luís Ceze, Doug Burger

As improvements in per-transistor speed and energy efficiency diminish, radical departures from conventional approaches are becoming critical to improving the performance and energy efficiency of general-purpose processors. We propose a solution—from circuit to compiler—that enables general-purpose use of limited-precision, analog hardware to accelerate “approximable” code—code that can tolerate imprecise execution. We utilize an algorithmic transformation that automatically converts approximabl

Electrical and Electronic EngineeringEngineering
10
Article|64 citations·2012
Locality-aware dynamic VM reconfiguration on MapReduce clouds
Jongse Park, Daewoo Lee, Bo-Kyeong Kim, Jaehyuk Huh, Seungryoul Maeng

Cloud computing based on system virtualization, has been expanding its services to distributed data-intensive platforms such as MapReduce and Hadoop. Such a distributed platform on clouds runs in a virtual cluster consisting of a number of virtual machines. In the virtual cluster, demands on computing resources for each node may fluctuate, due to data locality and task behavior. However, current cloud services use a static cluster configuration, fixing or manually adjusting the computing capabil

Information SystemsComputer Science
11
Article|49 citations·2021
SLO-Aware Inference Scheduler for Heterogeneous Processors in Edge Platforms
Wonik Seo, Sang-Hoon Cha, Yeonjae Kim, Jaehyuk Huh, Jongse Park
SJR Q2ACM Transactions on Architecture and Code OptimizationOA

With the proliferation of applications with machine learning (ML), the importance of edge platforms has been growing to process streaming sensor, data locally without resorting to remote servers. Such edge platforms are commonly equipped with heterogeneous computing processors such as GPU, DSP, and other accelerators, but their computational and energy budget are severely constrained compared to the data center servers. However, as an edge platform must perform the processing of multiple machine

Computer Networks and CommunicationsComputer Science
12
Article|44 citations·2015
Axilog: language support for approximate hardware design
Amir Yazdanbakhsh, Divya Mahajan, Bradley Thwaites, Jongse Park, Anandhavel Nagendrakumar, Sindhuja Sethuraman, Kartik Ramkrishnan, Nishanthi Ravindran, Rudra Jariwala, Abbas Rahimi, Hadi Esmaeilzadeh, Kia Bazargan
Design, Automation, and Test in Europe

Relaxing the traditional abstraction of near-perfect accuracy in hardware design can lead to significant gains in energy efficiency, area, and performance. To exploit this opportunity, there is a need for design abstractions that can systematically incorporate approximation in hardware design. We introduce Axilog, a set of language annotations, that provides the necessary syntax and semantics for approximate hardware design and reuse in Verilog. Axilog enables the designer to relax the accuracy

Electrical and Electronic EngineeringEngineering
13
Article|42 citations·2015
Axilog: Language Support for Approximate Hardware Design
Amir Yazdanbakhsh, Divya Mahajan, Bradley Thwaites, Jongse Park, Anandhavel Nagendrakumar, Sindhuja Sethuraman, Kartik Ramkrishnan, Nishanthi Ravindran, Rudra Jariwala, Abbas Rahimi, Hadi Esmaeilzadeh, Kia Bazargan
Design, Automation & Test in Europe Conference & Exhibition (DATE), 2015

Relaxing the traditional abstraction of “near-perfect” accuracy in hardware design can lead to significant gains in energy efficiency, area, and performance. To exploit this opportunity, there is a need for design abstractions that can systematically incorporate approximation in hardware design. We introduce Axilog, a set of language annotations, that provides the necessary syntax and semantics for approximate hardware design and reuse in Verilog. Axilog enables the designer to relax the accurac

Hardware and ArchitectureComputer Science
14
Article|41 citations·2014
Rollback-free value prediction with approximate loads
Bradley Thwaites, Gennady Pekhimenko, Hadi Esmaeilzadeh, Amir Yazdanbakhsh, Onur Mutlu, Jongse Park, Girish Mururu, Todd C. Mowry

This paper demonstrates how to utilize the inherent error resilience of a wide range of applications to mitigate the memory wall -- the discrepancy between core and memory speed. We define a new microarchitecturally-triggered approximation technique called rollback-free value prediction. This technique predicts the value of safe-to-approximate loads when they miss in the cache without tracking mispredictions or requiring costly recovery from misspeculations. This technique mitigates the memory w

Hardware and ArchitectureComputer Science
15
Article|41 citations·2017
Scale-out acceleration for machine learning
Jongse Park, Hardik Sharma, Divya Mahajan, Joon Kyung Kim, Preston Olds, Hadi Esmaeilzadeh

The growing scale and complexity of Machine Learning (ML) algorithms has resulted in prevalent use of distributed general-purpose systems. In a rather disjoint effort, the community is focusing mostly on high performance single-node accelerators for learning. This work bridges these two paradigms and offers CoSMIC, a full computing stack constituting language, compiler, system software, template architecture, and circuit generators, that enable programmable acceleration of learning at scale. CoS

Hardware and ArchitectureComputer Science

Research Areas

Hardware and ArchitectureComputer Vision and Pattern RecognitionElectrical and Electronic EngineeringArtificial IntelligenceComputer Networks and CommunicationsInformation Systems

JongSeok Parkの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。