Skip to main content
QUICK REVIEW

[論文レビュー] The shrinking human protein coding complement: are there now fewer than 20,000 genes?

Iakes Ezkurdia, David Juan|arXiv (Cornell University)|Dec 26, 2013
Machine Learning in Bioinformatics参考文献 53被引用数 8
ひとこと要約

本研究は、プロテオミクスデータと進化的保存性、ゲノム特徴を統合することで、ヒトのタンパク質コード遺伝子数を再評価し、ペプチド検出の欠如と保存性の低さにより2,001遺伝子が非コードである可能性が高いと特定した。著者らは、これらの仮想非コード遺伝子を除外することで、現在のタンパク質コード遺伝子アーカイブを20,000遺伝子未満に縮小することを提案する。

ABSTRACT

Determining the full complement of protein-coding genes is a key goal of genome annotation. The most powerful approach for confirming protein coding potential is the detection of cellular protein expression through peptide mass spectrometry experiments. Here we map the peptides detected in 7 large-scale proteomics studies to almost 60% of the protein coding genes in the GENCODE annotation the human genome. We find that conservation across vertebrate species and the age of the gene family are key indicators of whether a peptide will be detected in proteomics experiments. We find peptides for most highly conserved genes and for practically all genes that evolved before bilateria. At the same time there is almost no evidence of protein expression for genes that have appeared since primates, or for genes that do not have any protein-like features or cross-species conservation. We identify 19 non-protein-like features such as weak conservation, no protein features or ambiguous annotations in major databases that are indicators of low peptide detection rates. We use these features to describe a set of 2,001 genes that are potentially non-coding, and show that many of these genes behave more like non-coding genes than protein-coding genes. We detect peptides for just 3% of these genes. We suggest that many of these 2,001 genes do not code for proteins under normal circumstances and that they should not be included in the human protein coding gene catalogue. These potential non-coding genes will be revised as part of the ongoing human genome annotation effort.

研究の動機と目的

  • プロテオミクス的証拠と進化的保存性を用いて、ヒトのタンパク質コード遺伝子の構成を再評価すること。
  • 弱いか完全にペプチド検出がないために、タンパク質コード遺伝子として誤って分類されている遺伝子を特定すること。
  • クロススプライス保存性やタンパク質様モチーフの欠如といった、タンパク質コード特徴の欠落を示す遺伝子を除外することで、ヒトゲノムアノテーションを精緻化すること。
  • 真のタンパク質コード遺伝子と非コードまたは偽遺伝子トランスクリプトを区別するための体系的フレームワークを提供すること。
  • 通常の生物学的条件下でタンパク質コードの基準を満たさない遺伝子を特定することで、ヒトゲノムアノテーションの継続的取り組みを支援すること。

提案手法

  • GENCODEヒトゲノムアノテーションに、7つの大規模プロテオミクス研究からのペプチドをマッピングし、タンパク質発現を評価した。
  • 脊椎動物における進化的保存性と遺伝子ファミリーの年齢を、タンパク質コード可能性の指標として用いた。
  • 弱い保存性、タンパク質特徴の欠如、曖昧なデータベースアノテーションといった、ペプチド検出率が低いと関連する19の非タンパク質様特徴を同定した。
  • これらの特徴に基づいて遺伝子を分類し、ペプチド証拠が最小限の2,001個の候補非コード遺伝子リストを作成した。
  • これらの候補遺伝子の機能的および進化的挙動を、既知のタンパク質コード遺伝子と非コード遺伝子と比較して評価した。
  • 統計解析を用いて、ゲノム全体における保存性、遺伝子年齢、ペプチド検出頻度の相関を評価した。

実験結果

リサーチクエスチョン

  • RQ1最近のプロテオミクスおよび進化的データを踏まえると、ヒトゲノムにおけるタンパク質コード遺伝子数は20,000未満であると考えられるか?
  • RQ2質量分析実験におけるペプチド検出の低さまたは欠如を予測するゲノム特徴は何か?
  • RQ3進化的保存性と遺伝子ファミリーの年齢は、検出可能なタンパク質発現とどのように相関するか?
  • RQ4非タンパク質様特徴を示す遺伝子は、非コードRNAと比較してタンパク質コード遺伝子に類似している程度はどの程度か?
  • RQ5真のタンパク質コード遺伝子と非コードまたは偽遺伝子トランスクリプトを区別するための体系的フレームワークを開発可能か?

主な発見

  • 本研究で同定された2,001個の候補非コード遺伝子のうち、ペプチドが検出されたのはわずか3%にとどまった。
  • Bilaterian輻射以降に進化したが、クロススプライス保存性を欠く遺伝子は、ほとんどペプチド検出がなく、タンパク質コード可能性が低いことが示された。
  • 高い保存性を示す遺伝子および古代の遺伝子ファミリー(バイラテリア以前)に由来する遺伝子は、強いペプチド検出を示し、タンパク質コード状態を支持した。
  • 本研究では、弱い保存性、タンパク質ドメインの欠如、曖昧なデータベースアノテーションといった19の具体的なゲノム特徴が、ペプチド検出率の低さを強く示す指標であると特定した。
  • 現在のヒトタンパク質コード遺伝子アーカイブにおいて、これらの特徴を欠く遺伝子の大多数は、機能的なタンパク質コード遺伝子ではない可能性が高い。
  • 著者らは、ヒトタンパク質コード遺伝子数を下方修正すべきであると結論づけ、最終的な数は20,000遺伝子未満である可能性が高いと述べた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。