Skip to main content
QUICK REVIEW

[論文レビュー] An approach to describing and analysing bulk biological annotation quality: a case study using UniProtKB

M. J. Bell, CS Gillespie|Aug 10, 2012
Biomedical Text Mining and Ontologies被引用数 5
ひとこと要約

本論文は、UniProtKBのアノテーションにおける語の頻度分布をパワー則フィッティングとジプフの最小努力の原則を用いて分析することにより、バイオロジカルアノテーションの品質を評価する新規手法を提案する。手動アノテーション(Swiss-Prot)は当初、より高い品質(高いα値)を示すが、時間経過とともに品質が低下するのに対し、自動アノテーション(TrEMBL)は一貫して低いα値を示しており、読者の負担が増加し、アノテーターの利便性に寄り添う傾向にあることを示唆しており、αがアノテーション品質の妥当な代理指標である可能性を示している。

ABSTRACT

Motivation: Annotations are a key feature of many biological databases, used to convey our knowledge of a sequence to the reader. Ideally, annotations are curated manually, however manual curation is costly, time consuming and requires expert knowledge and training. Given these issues and the exponential increase of data, many databases implement automated annotation pipelines in an attempt to avoid un-annotated entries. Both manual and automated annotations vary in quality between databases and annotators, making assessment of annotation reliability problematic for users. The community lacks a generic measure for determining annotation quality and correctness, which we look at addressing within this article. Specifically we investigate word reuse within bulk textual annotations and relate this to Zipf's Principle of Least Effort. We use UniProt Knowledge Base (UniProtKB) as a case study to demonstrate this approach since it allows us to compare annotation change, both over time and between automated and manually curated annotations. Results: By applying power-law distributions to word reuse in annotation, we show clear trends in UniProtKB over time, which are consistent with existing studies of quality on free text English. Further, we show a clear distinction between manual and automated analysis and investigate cohorts of protein records as they mature. These results suggest that this approach holds distinct promise as a mechanism for judging annotation quality. Availability: Source code is available at the authors website: http://homepages.cs.ncl.ac.uk/m.j.bell1/annotation. Contact: phillip.lord@newcastle.ac.uk

研究の動機と目的

  • 外部メタデータやオントロジーに依存せずに、バイオロジカルアノテーションの品質を評価する汎用的でテキストのみに依存する指標を開発すること。
  • 語の頻度分布の時間的変化が、アノテーション品質やキュレーション実践の変化を反映しているかどうかを調査すること。
  • 言語的パターンとジプフの最小努力の原則を用いて、手動(Swiss-Prot)と自動(TrEMBL)アノテーションの品質を比較すること。
  • 語の頻度のパワー則フィッティングから得られるパラメータαが、アノテーション品質の信頼できる代理指標として機能するかどうかを評価すること。
  • この手法が、低品質または生物学的でないコンテンツをアノテーション内で検出する可能性を検討すること。

提案手法

  • 著者らは、1998年から2012年までのUniProtKBバージョンに含まれるすべての自由テキストアノテーションを抽出し、Swiss-ProtとTrEMBLに焦点を当てた。
  • すべてのアノテーションにおける語の頻度カウントを算出し、ランク付けされた語の頻度にパワー則分布をフィッティングし、スケーリングパラメータαを推定した。
  • αの値はジプフの最小努力の原則に基づいて解釈され、高いα値はより予測可能で読者にやさしい言語であることを示唆する。
  • 著者らは、時間経過に伴うα値の比較を通じて、アノテーション品質の傾向を検出するとともに、手動と自動アノテーションの間での差を分析した。
  • タンパク質エントリのコhortを成熟過程に沿って分析し、初期アノテーションから成熟アノテーションへの変化に伴うαの変化を評価した。
  • この手法はテキストコンテンツのみに依存するため、オントロジーまたは証拠コードの使用に関係なく、自由テキストアノテーションを備えた任意のデータベースに適用可能である。
Figure 1: Outline view of the data extraction process. (1) Initially we download a complete dataset for a given database version in flat file format. (2) We then extract the comment lines (lines beginning with ‘CC’, the comment indicator). (3) We remove comment blocks and properties (as defined in t
Figure 1: Outline view of the data extraction process. (1) Initially we download a complete dataset for a given database version in flat file format. (2) We then extract the comment lines (lines beginning with ‘CC’, the comment indicator). (3) We remove comment blocks and properties (as defined in t

実験結果

リサーチクエスチョン

  • RQ1バイオロジカルアノテーションにおける語の頻度のパワー則フィッティングが、アノテーション品質の信頼できる代理指標として機能するか。
  • RQ2手動でキュレートされたアノテーションと自動アノテーションの両方において、語の頻度分布から得られるαパラメータは、時間経過とともにどのように変化するか。
  • RQ3αの変化が、読者ではなくアノテーターの負担増加といった、アノテーション実践の変化をどの程度反映しているか。
  • RQ4この手法は、成熟したエントリや新規に追加されたエントリの品質低下を検出できるか。
  • RQ5αパラメータは、生物学的でない、または信号が弱いコンテンツを同定するのにも有用か。

主な発見

  • Swiss-ProtおよびTrEMBLの両方においてαパラメータが時間経過とともに低下しており、読者の負担が増加し、アノテーション品質が低下していることを示している。
  • Swiss-Protは当初、高いα値を示しており、高品質で読者に配慮したアノテーションであったが、データ量の増加とキュレーションの圧力が高まるにつれ品質が低下した。
  • TrEMBLはSwiss-Protと比較して一貫して低いα値を示しており、自動アノテーションが読者の理解を最適化するのではなく、アノテーターの利便性に最適化されていることを示唆している。
  • UniProtKB内の成熟したエントリでは、αが時間経過とともにゆっくりではあるが継続的に低下しており、データベースの拡大に伴い、既に確立されたアノテーションでさえ品質が低下している可能性を示している。
  • 新規に追加されたエントリに対しても、αは時間経過とともに低下しており、新規レコードに対しても品質が向上していないことが示された。
  • この手法は、コメント行内に含まれる著作権表記などの生物学的でないコンテンツを効果的に検出できた。これは、アーティファクト検出の可能性を示している。
Figure 2: Cumulative distributions of words for various Swiss-Prot and TrEMBL versions, shown with logarithmic scales. The size (number of words) is shown along the $X$ axis while the probability is shown on the $Y$ axis. A point on the graph represents the probability that a word will occur $x$ or
Figure 2: Cumulative distributions of words for various Swiss-Prot and TrEMBL versions, shown with logarithmic scales. The size (number of words) is shown along the $X$ axis while the probability is shown on the $Y$ axis. A point on the graph represents the probability that a word will occur $x$ or

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。