Hokkaido University · 컴퓨터과학
Timur Madzhidov 교수의 연구실은 반응 중심의 화학정보학, 특히 반응의 구조-반응성 관계(QRPR) 모델링을 핵심으로 합니다. 반응의 복잡성을 반영한 '응집 그래프 반응'(CGR) 기반의 소프트웨어 도구인 CGRtools를 개발하며, 반응의 반응성 예측, 원자 간 매핑 보정, 반응 조건 변화에 따른 속도 상수 예측 등 고도화된 반응 분석을 수행합니다. 특히 다양한 용매와 온도 조건에서의 기질 반응 속도를 정량적으로 예측하는 모델 개발에 주력하고 있으며, 교차검증 전략의 개선을 통해 보다 객관적인 예측 성능 평가 방법을 제안하고 있습니다.
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
CGRtools is an open-source Python library aimed to handle molecular and reaction information. It is the sole library developed so far which can process condensed graph of reaction (CGR) handling. CGR provides the possibility for advanced operations with reaction information and could be used for reaction descriptor calculation, structure-reactivity modeling, atom-to-atom mapping comparison and correction, reaction center extraction, reaction balancing, and some other related tasks. Unlike other
Nowadays, the problem of the model's applicability domain (AD) definition is an active research topic in chemoinformatics. Although many various AD definitions for the models predicting properties of molecules (Quantitative Structure-Activity/Property Relationship (QSAR/QSPR) models) were described in the literature, no one for chemical reactions (Quantitative Reaction-Property Relationships (QRPR)) has been reported to date. The point is that a chemical reaction is a much more complex object th
An approach for the prediction of rate constants of chemical reactions, based on the representation of a chemical reaction as a condensed graph, has been tested on more than 1000 bimolecular nucleophilic substitution reactions with neutral nucleophiles in 38 solvents. Molecular fragment descriptors, temperature, and solvent parameters characterizing solvation power have been used in the reaction modeling. The obtained models ensure a good correlation between the predicted and experimental values
Modern QSAR approaches have wide practical applications in drug discovery for designing potentially bioactive molecules. If such models are based on the use of 2D descriptors, important information contained in the spatial structures of molecules is lost. The major problem in constructing models using 3D descriptors is the choice of a putative bioactive conformation, which affects the predictive performance. The multi-instance (MI) learning approach considering multiple conformations in model tr
By means of a structural representation of the chemical reactivity as a condensed graph a model predicting rate constants of the bimolecular elimination reaction is derived for the first time. The model developed enables the prediction of rate constants of reactions proceeding in different solvents or water-organic mixtures at different temperatures. It demonstrates a good predictive performance: a mean square deviation of predicted values from experimental ones is less than 0.7 logarithmic unit
In this article, we consider cross-validation of the quantitative structure-property relationship models for reactions and show that the conventional k-fold cross-validation (CV) procedure gives an 'optimistically' biased assessment of prediction performance. To address this issue, we suggest two strategies of model cross-validation, 'transformation-out' CV, and 'solvent-out' CV. Unlike the conventional k-fold cross-validation approach that does not consider the nature of objects, the proposed p
The electronic structure of charge-transfer complexes of organoselenium compounds with diiodine has been studied at several levels of theory (Hartree-Fock, second order Møller-Plesset, and density functional theory). The complexation energies, optimized geometries, and the topology of the electron density and its Laplacian distribution, including domain averaged properties, have been analyzed. Special attention was paid to the influence of basis set superposition error on the energy of complexat
The most widely used QSAR approaches are mainly based on 2D molecular representation which ignores stereoconfiguration and conformational flexibility of compounds. 3D QSAR uses a single conformer of each compound which is difficult to choose reasonably. 4D QSAR uses multiple conformers to overcome the issues of 2D and 3D methods. However, many of existing 4D QSAR models suffer from the necessity to pre-align conformers, while alignment-independent approaches often ignore stereoconfiguration of c
By the structural representation of a chemical reaction in the form of a condensed graph a model allowing the prediction of rate constants (logk) of Diels–Alder reactions performed in different solvents and at different temperatures is constructed for the first time. The model demonstrates good agreement between the predicted and experimental logk values: the mean squared error is less than 0.75 log units. Erroneous predictions correspond to reactions in which reagents contain rarely occurring s
Here, we discuss a reaction standardization protocol followed by a comparison of popular Atom-to-atom mapping (AAM) tools (ChemAxon, Indigo, RDTool, NextMove and RXNMapper) as well as some consensus AAM strategies. For this purpose, a dataset of 1851 manually curated and mapped reactions was prepared (the Golden dataset) and used as a reference set. It has been found that RXNMapper possesses the highest accuracy, despite the fact that it has some clear disadvantages. Finally, RXNMapper was selec
The selection of experimental conditions leading to a reasonable yield is an important and essential element for the automated development of a synthesis plan and the subsequent synthesis of the target compound. The classical QSPR approach, requiring one-to-one correspondence between chemical structure and a target property, can be used for optimal reaction conditions prediction only on a limited scale when only one condition component (e.g., catalyst or solvent) is considered. However, a partic
The synthesis of the desired chemical compound is the main task of synthetic organic chemistry. The predictions of reaction conditions and some important quantitative characteristics of chemical reactions as yield and reaction rate can substantially help in the development of optimal synthetic routes and assessment of synthesis cost. Theoretical assessment of these parameters can be performed with the help of modern machine-learning approaches, which use available experimental data to develop pr
Pharmacophore modeling is usually considered as a special type of virtual screening without probabilistic nature. Correspondence of at least one conformation of a molecule to pharmacophore is considered as evidence of its bioactivity. We show that pharmacophores can be treated as one-class machine learning models, and the probability the reflecting model's confidence can be assigned to a pharmacophore on the basis of their precision of active compounds identification on a calibration set. Two sc