Jaehyung Kim
Yonsei University · Computer Science
About the Lab
Professor Jaehyung Kim's research lab specializes in machine learning and artificial intelligence, with a strong focus on addressing critical challenges in real-world data distributions. The lab investigates class imbalance in deep learning, developing innovative techniques such as data augmentation through cross-class translation and pseudo-label refinement to improve model generalization on minority classes. Additionally, the lab explores secure computation, particularly in fully homomorphic encryption, with a focus on optimizing bootstrapping mechanisms in the BFV scheme. Beyond AI, the lab also examines human well-being through the lens of environmental psychology, studying the relationships between restorative environments, leisure satisfaction, and psychological happiness in outdoor activities.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15In most real-world scenarios, labeled training datasets are highly class-imbalanced, where deep neural networks suffer from generalizing to a balanced testing criterion. In this paper, we explore a novel yet simple way to alleviate this issue by augmenting less-frequent classes via translating samples (e.g., images) from more-frequent classes. This simple approach enables a classifier to learn more generalizable features of minority classes, by transferring and leveraging the diversity of the ma
While semi-supervised learning (SSL) has proven to be a promising way for leveraging unlabeled data when labeled data is scarce, the existing SSL algorithms typically assume that training class distributions are balanced. However, these SSL algorithms trained under imbalanced class distributions can severely suffer when generalizing to a balanced testing criterion, since they utilize biased pseudo-labels of unlabeled data toward majority classes. To alleviate this issue, we formulate a convex op
In most real-world scenarios, labeled training datasets are highly class-imbalanced, where deep neural networks suffer from generalizing to a balanced testing criterion. In this paper, we explore a novel yet simple way to alleviate this issue by augmenting less-frequent classes via translating samples (e.g., images) from more-frequent classes. This simple approach enables a classifier to learn more generalizable features of minority classes, by transferring and leveraging the diversity of the ma
The purpose of this study was to investigate the relationship model of perceived restorative environment, leisure satisfaction, resilience and psychological happiness in university students’ outdoor sports participation. To achieve the goal of this study, a total 193 surveys collected from university in Seoul, Kyuggi, Chung-Chang, and Kangwoon areas were utilized for analyzing. frequency analysis, exploratory factor analysis, reliability analysis, confirmatory factor analysis and structural equa
Bootstrapping is currently the only known method for constructing fully homomorphic encryptions. In the BFV scheme specifically, bootstrapping aims to reduce the error of a ciphertext while preserving the encrypted plaintext. The existing BFV bootstrapping methods follow the same pipeline, relying on the evaluation of a digit extraction polynomial to annihilate the error located in the least significant digits. However, due to its strong dependence on performance, bootstrapping could only utiliz
The purpose of this study was to investigate the relationship among perceived restorative environment, place Attachment and psychological happiness for golf participants. To achieve the goal of this study, a total of 230 questionnaires were distributed and 230 copies were collected back. Out of those returned questionnaires, insincerely replied or double-replied questionnaires were excluded and finally 224 questionnaires were analyzed for this study. For analysis of the data, frequency analysis,
While semi-supervised learning (SSL) has proven to be a promising way for leveraging unlabeled data when labeled data is scarce, the existing SSL algorithms typically assume that training class distributions are balanced. However, these SSL algorithms trained under imbalanced class distributions can severely suffer when generalizing to a balanced testing criterion, since they utilize biased pseudo-labels of unlabeled data toward majority classes. To alleviate this issue, we formulate a convex op
이 연구는 한국 골프장산업의 현황을 분석하고 공급과잉을 진단함으로써 대안에 필요한 기초자료를 제시하고자 하였다. 이에 골프장산업과 관련된 각종 참고문헌, 통계자료, 전문기관들의 자료, 골프장운영 및 전문가들의 인터뷰를 통해 다음과 같은 단서들을 산출하였다. 즉 이미 공급과잉의 고위험군으로 제주권, 위험군으로는 강원권, 주의군으로는 충청권, 경상권, 전라권, 그리고 주의군 및 관찰군으로는 서울․경기권이 각각 분류되었다. 따라서 첫째, 현재 지방자치단체 주관으로 무차별적으로 추진되고 있는 골프장공급에 정부차원의 관리가 요구되고 있다. 둘째, 향후 살아남기 위한 경쟁과 국제경쟁력 또한 키울 수 있는 골프장 전문경영인이 요구되고 있다. 셋째, 공급과잉 해소를 위해 일회성 방문이 아닌 단골고객 유치와 지속적 방문이 성립될 수 있도록 생활체육활성화를 위한 골프장 공급구조와 객단가 위주의 운영시스템에서 가동률 위주의 영업전략으로 전환되어야 할 것이다.
The purpose of this study is to investigate the relationships between self-management ability, leisure engagement and active aging for the middle-aged and the elderly golf participants. To achieve the goal of this study, a total of 250 questionnaires were distributed and 250 copies were collected back. Out of those returned questionnaires, insincerely replied or double-replied questionnaires were excluded and finally 232 questionnaires were analyzed for this study. The data were analyzed by freq
Large language models (LLMs) have made significant advancements in various natural language processing tasks, including question answering (QA) tasks. While incorporating new information with the retrieval of relevant passages is a promising way to improve QA with LLMs, the existing methods often require additional fine-tuning which becomes infeasible with recent LLMs. Augmenting retrieved passages via prompting has the potential to address this limitation, but this direction has been limitedly
The purpose of this study was to investigate the relationship model of recreation specialization, resilience and psychological happiness in university students’ outdoor sports participation. To achieve the goal of this study, a total 289 surveys collected from university in Seoul, Kyuggi, Chung-Chang, and Kangwoon areas were utilized for analyzing. frequency analysis, exploratory factor analysis, reliability analysis, confirmatory factor analysis and structural equating modeling were conducted u
The Cheon–Kim–Kim–Song (CKKS) scheme is renowned for its efficiency in encrypted computing over real numbers. However, it lacks an important functionality that most exact schemes have, an efficient modular reduction. This derives from the fundamental difference in encoding structure. The CKKS scheme encodes messages to the least significant bits, while the other schemes encode to the most significant bits (or in an equivalent manner). As a result, CKKS could enjoy an efficient rescaling but lost
The success of NLP systems often relies on the availability of large, high-quality datasets. However, not all samples in these datasets are equally valuable for learning, as some may be redundant or noisy. Several methods for characterizing datasets based on model-driven meta-information (e.g., model’s confidence) have been developed, but the relationship and complementary effects of these methods have received less attention. In this paper, we introduce infoVerse, a universal framework for data
Research Areas
Dive deeper into Jaehyung Kim's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.