Skip to main content
QUICK REVIEW

[论文解读] Ab initio identification of putative human transcription factor binding sites by comparative genomics

Davide Corà, Carl Herrmann|ArXiv.org|May 3, 2005
Genomics and Chromatin Dynamics参考文献 33被引用 6
一句话总结

本研究提出一种比较基因组学方法,通过扫描人类与小鼠基因组中保守的上游区域,实现从头识别人类转录因子结合位点。通过检测在共保守基因集中富集的5–8 bp基序,并结合基因本体论(GO)富集分析和微阵列数据的共表达性筛选,该方法识别出已知及新型候选调控基序,证明其在无需预先知晓转录因子的情况下,可有效优先识别具有功能相关性的顺式调控元件。

ABSTRACT

We discuss a simple and powerful approach for the ab initio identification of cis-regulatory motifs involved in transcriptional regulation. The method we present integrates several elements: human-mouse comparison, statistical analysis of genomic sequences and the concept of coregulation. We apply it to a complete scan of the human genome. By using the catalogue of conserved upstream sequences collected in the CORG database we construct sets of genes sharing the same overrepresented motif (short DNA sequence) in their upstream regions both in human and in mouse. We perform this construction for all possible motifs from 5 to 8 nucleotides in length and then filter the resulting sets looking for two types of evidence of coregulation: first, we analyze the Gene Ontology annotation of the genes in the set, searching for statistically significant common annotations; second, we analyze the expression profiles of the genes in the set as measured by microarray experiments, searching for evidence of coexpression. The sets which pass one or both filters are conjectured to contain a significant fraction of coregulated genes, and the upstream motifs characterizing the sets are thus good candidates to be the binding sites of the TF's involved in such regulation. In this way we find various known motifs and also some new candidate binding sites.

研究动机与目标

  • 开发一种无需预先知晓调控因子的从头识别人类转录因子结合位点的方法。
  • 利用人类与小鼠基因组之间的进化保守性,优先识别具有功能相关性的顺式调控基序。
  • 通过共享上游基序识别共调控基因集,并利用功能注释和表达数据验证其生物一致性。
  • 通过整合多种生物学证据,将真实调控基序与随机序列模式区分开来。

提出的方法

  • 使用保守上游区域数据库(CORG)对人类与小鼠的上游序列进行全基因组比较。
  • 在保守的上游区域中扫描所有可能的5–8 bp DNA基序,检测其富集性。
  • 将两个物种中共享相同富集基序的基因归类为潜在的调控基因集。
  • 利用基因本体论(GO)富集分析筛选基序集合,检测其是否存在统计学上显著的共享生物功能。
  • 应用基于微阵列的共表达分析,检测基因集中基因的表达模式是否具有相关性。
  • 将通过至少一项筛选(GO富集或共表达)的基序优先列为高置信度候选转录因子结合位点。

实验结果

研究问题

  • RQ1在人类上游区域中,哪些短序列基序在人类与小鼠之间具有进化保守性,并在特定基因集中呈现富集?
  • RQ2共享相同保守上游基序的基因是否在基因本体论注释中表现出显著的功能相似性?
  • RQ3在微阵列实验中,共享相同上游基序的基因是否表现出共表达的证据?
  • RQ4结合GO注释与表达谱数据是否能提升对具有生物相关性的转录因子结合位点的识别能力?
  • RQ5哪些新颖或此前未被表征的基序被识别为高置信度的转录因子结合候选?

主要发现

  • 该方法成功识别出多个已知的转录因子结合基序,验证了其生物学相关性。
  • 大量基序集合在共享的基因本体论术语中表现出统计学上显著的富集,表明相关基因具有功能一致性。
  • 许多基序集合在微阵列数据中表现出共表达模式,支持其在共调控基因调控中的作用。
  • 整合保守性、富集性及功能证据,显著提高了候选基序选择的特异性,优于仅依赖序列保守性的方法。
  • 该方法识别出一组此前未与已知转录因子关联的新型候选基序,提示其为值得功能验证的新调控元件。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。