[论文解读] The shrinking human protein coding complement: are there now fewer than 20,000 genes?
本研究通过整合蛋白质组学数据、进化保守性及基因组特征,重新评估了人类蛋白质编码基因的数量,发现2,001个基因因缺乏肽段检测和保守性差而可能为非编码基因。作者建议通过排除这些疑似非编码基因,将当前的蛋白质编码基因目录减少至20,000个以下。
Determining the full complement of protein-coding genes is a key goal of genome annotation. The most powerful approach for confirming protein coding potential is the detection of cellular protein expression through peptide mass spectrometry experiments. Here we map the peptides detected in 7 large-scale proteomics studies to almost 60% of the protein coding genes in the GENCODE annotation the human genome. We find that conservation across vertebrate species and the age of the gene family are key indicators of whether a peptide will be detected in proteomics experiments. We find peptides for most highly conserved genes and for practically all genes that evolved before bilateria. At the same time there is almost no evidence of protein expression for genes that have appeared since primates, or for genes that do not have any protein-like features or cross-species conservation. We identify 19 non-protein-like features such as weak conservation, no protein features or ambiguous annotations in major databases that are indicators of low peptide detection rates. We use these features to describe a set of 2,001 genes that are potentially non-coding, and show that many of these genes behave more like non-coding genes than protein-coding genes. We detect peptides for just 3% of these genes. We suggest that many of these 2,001 genes do not code for proteins under normal circumstances and that they should not be included in the human protein coding gene catalogue. These potential non-coding genes will be revised as part of the ongoing human genome annotation effort.
研究动机与目标
- 利用蛋白质组学证据和进化保守性,重新评估人类蛋白质编码基因的组成。
- 识别因肽段检测微弱或缺失而可能被错误分类为蛋白质编码基因的基因。
- 通过排除缺乏跨物种保守性和蛋白质样基序等关键蛋白质编码特征的基因,优化人类基因组注释。
- 建立系统性框架,以区分真正的蛋白质编码基因与非编码或假基因转录本。
- 通过标记在正常生物学条件下不满足蛋白质编码标准的基因,支持人类基因组注释的持续改进工作。
提出的方法
- 将七项大规模蛋白质组学研究中鉴定的肽段比对至GENCODE人类基因组注释,以评估蛋白质表达情况。
- 利用脊椎动物间的进化保守性及基因家族年龄,作为蛋白质编码潜力的指标。
- 识别出19种非蛋白质样特征(如保守性弱、缺乏蛋白质特征、数据库注释模糊)与低肽段检测率相关。
- 基于这些特征对基因进行分类,生成一份包含2,001个候选非编码基因的列表,其肽段证据极少。
- 评估这些候选基因的功能与进化行为,将其与已知的蛋白质编码基因和非编码基因进行比较。
- 使用统计分析方法,关联全基因组范围内的保守性、基因年龄与肽段检测频率。
实验结果
研究问题
- RQ1在近期蛋白质组学与进化数据的支持下,人类基因组中蛋白质编码基因数量是否少于20,000个?
- RQ2哪些基因组特征可预测质谱实验中肽段检测率低或完全缺失?
- RQ3进化保守性与基因家族年龄如何与可检测的蛋白质表达相关联?
- RQ4具有非蛋白质样特征的基因在多大程度上表现得更像非编码RNA而非蛋白质编码基因?
- RQ5能否建立一个系统性框架,以区分真正的蛋白质编码基因与非编码或假基因转录本?
主要发现
- 在本研究识别出的2,001个候选非编码基因中,仅有3%检测到了肽段。
- 在两侧对称动物辐射后演化出且缺乏跨物种保守性的基因,几乎未检测到肽段,表明其蛋白质编码潜力极低。
- 高度保守的基因以及来自古老基因家族(原始两侧对称动物之前)的基因表现出强烈的肽段检测,支持其蛋白质编码状态。
- 本研究识别出19项特定的基因组特征(如保守性弱、缺乏蛋白质结构域、数据库注释模糊)是低肽段检测的强预测指标。
- 当前人类蛋白质编码基因目录中,若缺乏这些特征的基因,极有可能并非功能性蛋白质编码基因。
- 作者得出结论:人类蛋白质编码基因数量应下调修订,最终数量很可能低于20,000个。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。