Yeongjae Choi
Korea Advanced Institute of Science and Technology · Biochemistry, Genetics and Molecular Biology
About the Lab
Professor Yeongjae Choi's research lab specializes in high-performance computing and bioinformatics, focusing on optimizing data management for scientific workloads and advancing structural biology through computational methods. The lab develops innovative algorithms and web-based resources for declustering multidimensional datasets in parallel databases, improving I/O efficiency in high-performance computing environments. It also creates tools like PSIMAP and SNP@Domain to map protein interactions and annotate disease-associated genetic variations within protein domains, integrating structural and sequence-based domain information for biological insight. The lab bridges computer science and life sciences by designing scalable, visualization-driven solutions for complex biological and computational data.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
4UNLABELLED: Protein Structural Interactome map (PSIMAP) is a global interaction map that describes domain-domain and protein-protein interaction information for known Protein Data Bank structures. It calculates the Euclidean distance to determine interactions between possible pairs of structural domains in proteins. PSIbase is a database and file server for protein structural interaction information calculated by the PSIMAP algorithm. PSIbase also provides an easy-to-use protein domain assignmen
The single nucleotide polymorphisms (SNPs) in conserved protein regions have been thought to be strong candidates that alter protein functions. Thus, we have developed SNP@Domain, a web resource, to identify SNPs within human protein domains. We annotated SNPs from dbSNP with protein structure-based as well as sequence-based domains: (i) structure-based using SCOP and (ii) sequence-based using Pfam to avoid conflicts from two domain assignment methodologies. Users can investigate SNPs within pro
Declustering is a well known technique to achieve high performance for queries on parallel databases. We propose novel General Disk Module (GDM) based declustering algorithms, GDM Cartesian and GDM Circle, for distributing uniformly distributed multidimensional datasets to parallel disks, for datasets of any dimension. We compare the performance of the new approaches with several existing declustering algorithms, using variable numbers of disks, and with variable shapes and dimensions of the dat
Large multidimensional arrays are a common data type in high-performance scientific applications. Without special techniques for handling access to these arrays, I/O can easily become a large fraction of execution time for applications using these arrays, especially on parallel platforms. We show how to reduce the parallel I/O bottleneck for array data in closely-synchronized SPMD applications on distributed-memory platforms, through the use of server-directed I/O. This method allows array data
Research Areas
Dive deeper into Yeongjae Choi's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.