Sung‐Hoon Kim
Korea Advanced Institute of Science and Technology · Computer Science
About the Lab
Professor Sung-Hoon Kim's research lab specializes in computational science and software engineering, with a focus on improving software quality and reliability through machine learning and data-driven analysis. The lab develops innovative techniques for latent bug detection using change classification and defect prediction models trained on software repository data, while also addressing challenges in simulation setup and free energy calculations for molecular systems. Their work bridges software engineering and computational chemistry, particularly through tools like CHARMM-GUI for ligand modeling and alchemical free energy simulations. The lab emphasizes practical, scalable solutions for both software development and molecular simulation workflows.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15This paper introduces a new technique for finding latent software bugs called change classification. Change classification uses a machine learning classifier to determine whether a new software change is more similar to prior buggy changes, or clean changes. In this manner, change classification predicts the existence of bugs in software changes. The classifier is trained using features (in the machine learning sense) extracted from the revision history of a software project, as stored in its so
Reading ligand structures into any simulation program is often nontrivial and time consuming, especially when the force field parameters and/or structure files of the corresponding molecules are not available. To address this problem, we have developed Ligand Reader & Modeler in CHARMM-GUI. Users can upload ligand structure information in various forms (using PDB ID, ligand ID, SMILES, MOL/MOL2/SDF file, or PDB/mmCIF file), and the uploaded structure is displayed on a sketchpad for verification
Many software defect prediction models have been built using historical defect data obtained by mining software repositories (MSR). Recent studies have discovered that data so collected contain noises because current defect collection practices are based on optional bug fix keywords or bug report links in change logs. Automatically collected defect data based on the change logs could include noises.
The number of bugs (or fixes) is a common factor used to measure the quality of software and assist bug related analysis. For example, if software files have many bugs, they may be unstable. In comparison, the bug-fix time - the time to fix a bug after the bug was introduced - is neglected. We believe that the bug-fix time is an important factor for bug related analysis, such as measuring software quality. For example, if bugs in a file take a relatively long time to be fixed, the file may have
We have performed scanning tunneling microscopy and differential tunneling conductance (dI/dV) mapping for the surface of the three-dimensional topological insulator Bi(2)Se(3). The fast Fourier transformation applied to the dI/dV image shows an electron interference pattern near Dirac node despite the general belief that the backscattering is well suppressed in the bulk energy gap region. The comparison of the present experimental result with theoretical surface and bulk band structures shows t
Alchemical free energy simulations have long been utilized to predict free energy changes for binding affinity and solubility of small molecules. However, while the theoretical foundation of these methods is well established, seamlessly handling many of the practical aspects regarding the preparation of the different thermodynamic end states of complex molecular systems and the numerous processing scripts often remains a burden for successful applications. In this work, we present CHARMM-GUI <i>
This article provides technical descriptions of five fixed parameter calibration (FPC) methods, which were based on marginal maximum likelihood estimation via the EM algorithm, and evaluates them through simulation. The five FPC methods described are distinguished from each other by how many times they update the prior ability distribution and by how many EM cycles they use. Specifically, the five FPC methods included no prior weights updating and one EM cycle (NWU‐OEM) or multiple EM cycles (NW
It is a common understanding that identifying the same entity such as module, file, and function between revisions is important for software evolution related analysis. Most software evolution researchers use entity names, such as file names and function names, as entity identifiers based on the assumption that each entity is uniquely identifiable by its name. Unfortunately names change over time. In this paper, we propose an automated algorithm that identifies entity mapping at the function lev
Automatic bug finding tools tend to have high false positive rates: most warnings do not indicate real bugs. Usually bug finding tools prioritize each warning category. For example, the priority of "overflow " is 1 and the priority of "jumbled incremental" is 3, but the tools 'prioritization is not very effective. In this paper, we prioritize warning categories by analyzing the software change history. The underlying intuition is that if warnings from a category are resolved quickly by developer
This article extends the Bonett (2003a) approach to testing the equality of alpha coefficients from two independent samples to the case of m ≥ 2 independent samples. The extended Fisher‐Bonett test and its competitor, the Hakstian‐Whalen (1976) test, are illustrated with numerical examples of both hypothesis testing and power calculation. Computer simulations are used to compare the performance of the two tests and the Feldt (1969) test (for m = 2) in terms of power and Type I error control. It
Electron scattering in the topological surface state (TSS) of the topological insulator Bi1.5Sb0.5Te1.7Se1.3 was studied using quasiparticle interference observed by scanning tunneling microscopy. It was found that not only the 180° backscattering but also a wide range of backscattering angles of 100°-180° are effectively prohibited in the TSS. This conclusion was obtained by comparing the observed scattering vectors with the diameters of the constant-energy contours of the TSS, which were measu
Under item response theory (IRT), linking proficiency scales from separate calibrations of multiple forms of a test to achieve a common scale is required in many applications. Four IRT linking methods including the mean/mean, mean/sigma, Haebara, and Stocking‐Lord methods have been presented for use with single‐format tests. This study extends the four linking methods to a mixture of unidimensional IRT models for mixed‐format tests. Each linking method extended is intended to handle mixed‐format
Assuming item parameters on a test are known constants, the reliability coefficient for item response theory (IRT) ability estimates is defined for a population of examinees in two different ways: as (a) the product-moment correlation between ability estimates on two parallel forms of a test and (b) the squared correlation between the true abilities and estimates. Due to the bias of IRT ability estimates, the parallel-forms reliability coefficient is not generally equal to the squared-correlatio
Research Areas
Dive deeper into Sung‐Hoon Kim's research on Nubint
Open this lab's papers in the app to read with AI, summarize, and cite in your writing.