[论文解读] Computational identification of transcription factor binding sites by functional analysis of sets of genes sharing overrepresented upstream motifs
本文提出一种计算方法,通过检测上游基因区域中富集的基序并结合基因本体论(GO)术语的功能富集验证,识别转录因子结合位点。通过将具有富集基序的基因分组并检测显著的GO术语富集,该方法在 *S. cerevisiae* 中识别出已知和新型的调控元件,且通过经验随机化控制假发现率,相较于仅依赖表达数据的方法更具可靠性。
BACKGROUND: Transcriptional regulation is a key mechanism in the functioning of the cell, and is mostly effected through transcription factors binding to specific recognition motifs located upstream of the coding region of the regulated gene. The computational identification of such motifs is made easier by the fact that they often appear several times in the upstream region of the regulated genes, so that the number of occurrences of relevant motifs is often significantly larger than expected by pure chance. RESULTS: To exploit this fact, we construct sets of genes characterized by the statistical overrepresentation of a certain motif in their upstream regions. Then we study the functional characterization of these sets by analyzing their annotation to Gene Ontology terms. For the sets showing a statistically significant specific functional characterization, we conjecture that the upstream motif characterizing the set is a binding site for a transcription factor involved in the regulation of the genes in the set. CONCLUSIONS: The method we propose is able to identify many known binding sites in S. cerevisiae and new candidate targets of regulation by known transcription factors. Its application to less well studied organisms is likely to be valuable in the exploration of their regulatory interaction network.
研究动机与目标
- 通过利用功能注释而非仅依赖表达数据,改进转录因子结合位点的计算识别方法。
- 解决基于微阵列验证的局限性,如实验偏差以及对小基因集的低敏感性。
- 开发一种稳健的方法,利用基因本体论(GO)术语富集作为验证标准,检测具有生物意义的调控基序。
- 通过使用随机基因集的经验估计,控制基序-基因关联中的假阳性率。
- 通过提供独立于表达数据的功能验证框架,将该方法推广至研究较少的生物体。
提出的方法
- 针对所有5至8个核苷酸的基序,将上游区域中基序出现频率显著高于背景频率的基因归入特定基序的基因集。
- 通过超几何检验测试每个基序-基因集在GO术语中的显著富集,实现功能表征。
- 通过生成350万个典型大小(20个基因)的随机基因集,经验估计假发现率(FDR),并计算各GO分支中最佳P值的分布。
- FDR通过随机集产生的预期假发现数与观察到的显著关联数的比值计算得出,从而实现统计可信度的阈值选择。
- 该方法采用两步验证:首先检测基序富集,然后通过GO富集检测功能一致性,避免依赖微阵列数据。
- 该方法应用于 *S. cerevisiae*,结果与文献中已知的转录因子-基因相互作用(例如,参考文献[5])进行比较。
实验结果
研究问题
- RQ1基因本体论术语的功能富集能否作为计算识别的转录因子结合基序的可靠验证手段?
- RQ2在同时测试数千个基序和GO术语时,如何控制假阳性基序-基因关联?
- RQ3基于GO的验证是否能检测到因实验限制而被基于表达的方法遗漏的调控基序?
- RQ4该方法在注释较为充分的生物体(如 *S. cerevisiae*)中,能在多大程度上识别已知和新型的调控元件?
- RQ5当表达数据有限或不可用时,该方法是否可推广至研究较少的生物体?
主要发现
- 该方法通过基序关联基因集中GO术语的显著功能富集,成功识别出 *S. cerevisiae* 中许多已知的转录因子结合位点。
- 通过将富集基序与一致的生物学功能关联,该方法检测到已知转录因子的新型候选调控靶标。
- 通过经验随机化,该方法将假发现率控制在0.01,确保了识别出的基序-基因关联具有高度可信度。
- 基于GO的验证补充并扩展了基于表达的方法,尤其适用于信号不强的小型或功能一致的基因集。
- 该方法在不同基因集大小下均表现出稳健性,模拟结果显示无论选择何种基因集大小,FDR估计均保持一致。
- 与实验确定的转录因子靶标(参考文献[5])进行系统比较表明,具有显著GO富集的基序-基因集与已知TF调控基因的重叠高度显著(许多情况下P < 10⁻⁵)。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。