Skip to main content
QUICK REVIEW

[Paper Review] Prediction of a Gene Regulatory Network from Gene Expression Profiles With Linear Regression and Pearson Correlation Coefficient

Mehedi Hassan Onik, Shakhawat Ahmmed Nobin|arXiv (Cornell University)|May 2, 2018
Gene expression and cancer classification19 references3 citations
TL;DR

This study proposes a novel machine learning approach to reconstruct a cancer-specific gene regulatory network (GRN) using gene expression profiles. It applies linear regression to identify differentially expressed genes, followed by Pearson correlation to infer regulatory relationships, successfully identifying hub genes for potential cancer diagnostic targets, validated against biological databases.

ABSTRACT

Reconstruction of gene regulatory networks is the process of identifying gene dependency from gene expression profile through some computation techniques. In our human body, though all cells pose similar genetic material but the activation state may vary. This variation in the activation of genes helps researchers to understand more about the function of the cells. Researchers get insight about diseases like mental illness, infectious disease, cancer disease and heart disease from microarray technology, etc. In this study, a cancer-specific gene regulatory network has been constructed using a simple and novel machine learning approach. In First Step, linear regression algorithm provided us the significant genes those expressed themselves differently. Next, regulatory relationships between the identified genes has been computed using Pearson correlation coefficient. Finally, the obtained results have been validated with the available databases and literatures. We can identify the hub genes and can be targeted for the cancer diagnosis.

Motivation & Objective

  • To develop a simple and effective method for reconstructing gene regulatory networks from gene expression profiles.
  • To identify differentially expressed genes in cancer using linear regression for improved network specificity.
  • To infer regulatory interactions between genes using Pearson correlation coefficient.
  • To validate the predicted network against existing biological databases and literature.
  • To identify hub genes with potential as diagnostic or therapeutic targets in cancer.

Proposed method

  • Linear regression is applied to gene expression profiles to identify genes significantly differentially expressed across conditions.
  • The most significant genes from linear regression are selected as candidate regulators and targets for network construction.
  • Pearson correlation coefficient is computed between pairs of selected genes to quantify the strength and direction of linear relationships.
  • A regulatory relationship is inferred if the correlation exceeds a predefined threshold, indicating potential regulatory influence.
  • The resulting network is visualized and cross-validated using known biological interactions from public databases.
  • Hub genes are identified based on their high connectivity in the reconstructed network.

Experimental results

Research questions

  • RQ1Which genes are significantly differentially expressed in the cancer dataset, as identified by linear regression?
  • RQ2What regulatory relationships exist between differentially expressed genes, as measured by Pearson correlation?
  • RQ3Can the predicted gene regulatory network be validated using existing biological databases?
  • RQ4Which genes emerge as hubs in the reconstructed network, indicating potential regulatory importance?
  • RQ5Are the predicted regulatory interactions biologically plausible and consistent with known cancer-related pathways?

Key findings

  • Linear regression successfully identified a set of differentially expressed genes from the gene expression dataset, forming the basis for network construction.
  • Pearson correlation coefficient effectively captured significant linear relationships between gene pairs, enabling inference of regulatory interactions.
  • The reconstructed gene regulatory network showed consistency with known biological interactions when cross-validated against public databases.
  • Hub genes with high connectivity were identified, suggesting their potential role in regulating key cancer-related processes.
  • The method produced a biologically plausible network that highlights candidate genes for further experimental validation in cancer diagnosis.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.