Yonsei University · Computer Science
Professor Jiashun Mao's research lab specializes in the integration of machine learning, molecular simulation, and chemical informatics to advance computational drug discovery and materials science. The lab focuses on developing data-driven models that bridge molecular representation (such as IUPAC nomenclature and SMILES) with deep generative models, particularly diffusion models, for intelligent molecular design. A key research direction involves leveraging natural language processing techniques for chemical language to enable interpretable and editable molecular generation, while also improving the accuracy of property prediction—such as dielectric constants—for functional materials. The lab emphasizes the synergy between wet-lab experiments, molecular dynamics simulations, and AI-driven modeling to ensure physical realism and practical applicability.
Figures are computed from collected data and may differ slightly.
Early quantitative structure-activity relationship (QSAR) technologies have unsatisfactory versatility and accuracy in fields such as drug discovery because they are based on traditional machine learning and interpretive expert features. The development of Big Data and deep learning technologies significantly improve the processing of unstructured data and unleash the great potential of QSAR. Here we discuss the integration of wet experiments (which provide experimental data and reliable verific
Since the Simplified Molecular Input Line Entry System (SMILES) is oriented to the atomic-level representation of molecules and is not friendly in terms of human readability and editable, however, IUPAC is the closest to natural language and is very friendly in terms of human-oriented readability and performing molecular editing, we can manipulate IUPAC to generate corresponding new molecules and produce programming-friendly molecular forms of SMILES. In addition, antiviral drug design, especial
The IUPAC (International Union of Pure and Applied Chemistry) nomenclature is a globally recognized unique naming system which assigns names to chemical compounds. As a form of molecular representation closest to natural language, it allows to estimate molecular data in a large-scale pre-trained paradigm by employing machine learning approaches for natural language processing (NLP). Although, SMILES is currently popular molecular representation used by most generative models, different molecular
Recently, diffusion models have emerged as a promising paradigm for molecular design and optimization. However, most diffusion-based molecular generative models focus on modeling 2D graphs or 3D geometries, with limited research on molecular sequence diffusion models. The International Union of Pure and Applied Chemistry (IUPAC) names are more akin to chemical natural language than the Simplified Molecular Input Line Entry System (SMILES) for organic compounds. In this work, we apply an IUPAC-gu
Dielectric constant (DC, ε) is a fundamental parameter in material sciences to measure polarizability of the system. In industrial processes, its value is an imperative indicator, which demonstrates the dielectric property of material and compiles information including separation information, chemical equilibrium, chemical reactivity analysis, and solubility modeling. Since, the available ε-prediction models are fairly primitive and frequently suffer from serious failures especially when deals w
Protein-protein interactions are the basis of many protein functions, and understanding the contact and conformational changes of protein-protein interactions is crucial for linking protein structure to biological function. Although difficult to detect experimentally, molecular dynamics (MD) simulations are widely used to study the conformational ensembles and dynamics of protein-protein complexes, but there are significant limitations in sampling efficiency and computational costs. In this stud
ABSTRACT Obtaining positive and negative samples to examining several multifaceted brain diseases in clinical trials face significant challenges. We propose an innovative approach known as Adaptive Conditional Graph Diffusion Convolution (ACGDC) model. This model is tailored for the fusion of single cell multi-omics data and the creation of novel samples. ACGDC customizes a new array of edge relationship categories to merge single cell sequencing data and pertinent meta-information gleaned from
The electrical conductivity of block copolymer nanocomposites is governed by a complex interplay between nanofiller organization and block copolymer morphology. However, establishing a quantitative, predictive link between the molecular-scale structure and macroscopic properties remains a fundamental challenge. Here, we demonstrate that the local particle density field serves as a pivotal order parameter controlling the conductive network. By integrating hybrid particle-field molecular dynamics
Open papers in the app to read, cite, and organize with AI.