[Paper Review] A robust kernel machine regression towards biomarker selection in multi-omics datasets of osteoporosis for drug discovery
This paper proposes RobKMR, a robust kernel machine regression method that integrates multi-omics data to identify stable, biologically relevant biomarkers in osteoporosis despite data contamination and outliers. By leveraging a robust M-estimator-based kernel framework and a novel score test, RobKMR identifies significant gene triplets (e.g., DKK1, SMTN, DRGX) and top genes (DKK1, MTND5, FASTKD2, SIDT1) linked to bone mineral density, enabling effective drug repurposing with Tacrolimus, Ibandronate, Alendronate, and Bazedoxifene.
Many statistical machine approaches could ultimately highlight novel features of the etiology of complex diseases by analyzing multi-omics data. However, they are sensitive to some deviations in distribution when the observed samples are potentially contaminated with adversarial corrupted outliers (e.g., a fictional data distribution). Likewise, statistical advances lag in supporting comprehensive data-driven analyses of complex multi-omics data integration. We propose a novel non-linear M-estimator-based approach, "robust kernel machine regression (RobKMR)," to improve the robustness of statistical machine regression and the diversity of fictional data to examine the higher-order composite effect of multi-omics datasets. We address a robust kernel-centered Gram matrix to estimate the model parameters accurately. We also propose a robust score test to assess the marginal and joint Hadamard product of features from multi-omics data. We apply our proposed approach to a multi-omics dataset of osteoporosis (OP) from Caucasian females. Experiments demonstrate that the proposed approach effectively identifies the inter-related risk factors of OP. With solid evidence (p-value = 0.00001), biological validations, network-based analysis, causal inference, and drug repurposing, the selected three triplets ((DKK1, SMTN, DRGX), (MTND5, FASTKD2, CSMD3), (MTND5, COG3, CSMD3)) are significant biomarkers and directly relate to BMD. Overall, the top three selected genes (DKK1, MTND5, FASTKD2) and one gene (SIDT1 at p-value= 0.001) significantly bond with four drugs- Tacrolimus, Ibandronate, Alendronate, and Bazedoxifene out of 30 candidates for drug repurposing in OP. Further, the proposed approach can be applied to any disease model where multi-omics datasets are available.
Motivation & Objective
- Address the challenge of statistical sensitivity to outliers and non-normal distributions in multi-omics data analysis for complex diseases.
- Overcome limitations of traditional kernel machine methods that assume ideal data distributions and fail under adversarial or corrupted data.
- Develop a robust, non-linear statistical framework to detect higher-order composite effects across genomics, transcriptomics, and epigenomics data.
- Enable reliable biomarker selection and drug repurposing in osteoporosis using integrated multi-omics datasets with enhanced robustness.
- Provide a generalizable method applicable to any disease model with available multi-omics data.
Proposed method
- Propose a robust kernel machine regression (RobKMR) framework using M-estimators to minimize the influence of outliers and non-normal data distributions.
- Construct a robust kernel-centered Gram matrix to improve parameter estimation accuracy under data contamination.
- Introduce a robust score test to assess marginal and joint effects of features via Hadamard product of multi-omics data components.
- Apply the method to a real multi-omics dataset from Caucasian females with osteoporosis to identify inter-related risk factors.
- Perform molecular docking between 37 selected proteins and 35 candidate drugs to predict binding affinities and identify top drug candidates.
- Use 3D structural modeling and interaction profiling to validate protein-ligand interactions for top hits (e.g., hydrogen bonds, hydrophobic interactions).
Experimental results
Research questions
- RQ1Can a robust kernel machine regression method effectively detect higher-order composite effects in multi-omics data under non-ideal distributional assumptions?
- RQ2How does RobKMR perform in identifying biologically relevant biomarkers when data contain adversarial outliers or fictional distributions?
- RQ3Which multi-omics gene combinations show significant association with bone mineral density (BMD) in osteoporosis?
- RQ4Can the identified biomarkers be linked to existing drugs through molecular docking and binding affinity analysis?
- RQ5What is the predictive power and robustness of RobKMR compared to state-of-the-art methods in both simulated and real-world multi-omics datasets?
Key findings
- RobKMR identified three significant gene triplets—(DKK1, SMTN, DRGX), (MTND5, FASTKD2, CSMD3), and (MTND5, COG3, CSMD3)—with p-values ≤ 0.00001, all directly related to bone mineral density (BMD).
- The top four genes—DKK1, MTND5, FASTKD2, and SIDT1—showed strong statistical significance (p ≤ 0.00001 for DKK1, MTND5, FASTKD2; p ≤ 0.001 for SIDT1) and biological relevance to osteoporosis.
- Molecular docking identified four lead drugs—Tacrolimus, Ibandronate, Alendronate, and Bazedoxifene—with average binding affinity scores ≥ -7.5 kcal/mol against the top four target proteins.
- The DKK1_Alendronate complex formed six hydrogen bonds and significant hydrophobic interactions, indicating strong binding stability.
- The MTND5_Ibandronate complex formed five hydrogen bonds and key hydrophobic interactions, supporting high-affinity binding.
- Network and causal inference analyses confirmed that the selected genes are functionally interconnected and directly influence BMD, validating their biological relevance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.