Skip to main content
QUICK REVIEW

[论文解读] Novel Method for Mutational Disease Prediction using Bioinformatics Techniques and Backpropagation Algorithm

Ayad Ghany Ismaeel, Anar Auda Ablahad|arXiv (Cornell University)|Mar 3, 2013
Genetics, Bioinformatics, and Biomedical Research参考文献 9被引用 9
一句话总结

本文提出了一种新颖的生物信息学方法,结合序列比对工具(FASTA、CLUSTALW)与反向传播神经网络,用于预测突变性疾病,特别是乳腺癌。通过将患者蛋白序列与参考疾病相关蛋白(如BRCA1/BRCA2)进行比较,将突变分类为恶性,训练过程中均方误差达到0.0000001,成功实现预测。

ABSTRACT

Cancer is one of the most feared diseases in the world it has increased disturbingly and breast cancer occurs in one out of eight women, the prediction of malignancies plays essential roles not only in revealing human genome, but also in discovering effective prevention and treatment of cancers. Generally cancer disease driven by somatic mutations in an individual DNA sequence, or genome that accumulates during the lifetime of person. This paper is proposed a novel method can predict the disease by mutations despite The presence in gene sequence is not necessary it are malignant, so will be compare the protein of patient with the gene's protein of disease if there is difference between these two proteins then can say there is malignant mutations. This method will use bioinformatics techniques like FASTA, CLUSTALW, etc which shows whether malignant mutations or not, then training the backpropagation algorithm using all expected malignant mutations for a certain genes (e.g. BRCA1 and BRCA2) of disease, and using it to test whether patient is holder the disease or not. Implementing this novel method as the first way to predict the disease based on mutations in the sequence of the gene that causes the disease shows two decisions are achieved successfully, the first diagnose whether the patient has mutations of cancer or not using bioinformatics techniques the second classifying these mutations are related to breast cancer (e.g. BRCA1 and BRCA2) using backpropagation with mean square rate 0.0000001. Keywords-Gene sequence; Protein; Deoxyribonucleic Acid DNA; Malignant mutation; Bioinformatics; Back-propagation algorithm.

研究动机与目标

  • 开发一种基于基因突变而非仅基因存在与否的新方法,用于预测突变性疾病。
  • 通过将患者蛋白序列与已知疾病相关蛋白(如BRCA1、BRCA2)进行比较,识别恶性突变。
  • 应用生物信息学工具(FASTA、CLUSTALW)检测指示恶性的序列差异。
  • 利用已知的恶性突变数据对反向传播神经网络进行训练,以实现分类。
  • 利用机器学习实现对癌症相关突变的准确诊断与分类。

提出的方法

  • 利用FASTA与CLUSTALW对患者蛋白序列与参考蛋白序列进行序列比对。
  • 将患者蛋白序列与野生型疾病相关蛋白(如BRCA1、BRCA2)进行比较,以检测差异。
  • 基于结构与序列变异识别潜在的恶性突变。
  • 使用来自BRCA1与BRCA2基因的已知恶性突变数据对反向传播神经网络进行训练。
  • 在网络训练过程中采用0.0000001的均方误差率以实现收敛。
  • 利用训练完成的网络将新患者的序列分类为携带或不携带与疾病相关的突变。

实验结果

研究问题

  • RQ1患者与参考基因之间的蛋白序列差异是否能可靠指示恶性突变?
  • RQ2生物信息学工具(如FASTA与CLUSTALW)是否能在不依赖基因存在的前提下检测出具有生物学意义的突变?
  • RQ3反向传播神经网络是否能基于训练数据有效分类突变为致病性或良性?
  • RQ4所提出的方法是否能够有效区分BRCA1/BRCA2相关突变与非致病性变异?
  • RQ5该方法是否能仅依赖少量输入数据,以高精度预测突变性疾病?

主要发现

  • 该方法成功诊断出患者是否携带与癌症相关的突变,基于生物信息学序列比对。
  • 反向传播神经网络实现了稳定的训练,均方误差为0.0000001。
  • 患者与参考蛋白之间的序列差异被用作恶性突变的指示指标。
  • 该方法在分类与BRCA1和BRCA2基因相关的突变方面展示了可行性。
  • 序列比对工具与神经网络的整合实现了双阶段预测:检测与分类。
  • 该方法提供了一种新颖的、不依赖于基因存在的突变性疾病预测方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。