Skip to main content
QUICK REVIEW

[논문 리뷰] An approach to describing and analysing bulk biological annotation quality: a case study using UniProtKB

M. J. Bell, CS Gillespie|2012. 08. 10.
Biomedical Text Mining and Ontologies인용 수 5
한 줄 요약

이 논문은 유니프로트KB의 애너테이션에서 단어 빈도 분포를 분석함으로써 대량 생물학적 애너테이션 품질을 평가하는 새로운 방법을 제안한다. 이 방법은 힘의 법칙 피팅과 지프의 최소 노력 원리를 사용한다. 연구 결과, 수작업 애너테이션(Swiss-Prot)은 초반에 더 높은 품질(더 높은 α 값)을 보였지만 시간이 지남에 따라 품질이 떨어지며, 자동 애너테이션(TrEMBL)은 일관되게 낮은 α 값을 보여 독자가 더 많은 노력을 기울여야 하며, 이는 애너테이터의 편의성 향상으로 이어지는 경향을 보이며, α가 애너테이션 품질의 타당한 대체 지표가 될 수 있음을 시사한다.

ABSTRACT

Motivation: Annotations are a key feature of many biological databases, used to convey our knowledge of a sequence to the reader. Ideally, annotations are curated manually, however manual curation is costly, time consuming and requires expert knowledge and training. Given these issues and the exponential increase of data, many databases implement automated annotation pipelines in an attempt to avoid un-annotated entries. Both manual and automated annotations vary in quality between databases and annotators, making assessment of annotation reliability problematic for users. The community lacks a generic measure for determining annotation quality and correctness, which we look at addressing within this article. Specifically we investigate word reuse within bulk textual annotations and relate this to Zipf's Principle of Least Effort. We use UniProt Knowledge Base (UniProtKB) as a case study to demonstrate this approach since it allows us to compare annotation change, both over time and between automated and manually curated annotations. Results: By applying power-law distributions to word reuse in annotation, we show clear trends in UniProtKB over time, which are consistent with existing studies of quality on free text English. Further, we show a clear distinction between manual and automated analysis and investigate cohorts of protein records as they mature. These results suggest that this approach holds distinct promise as a mechanism for judging annotation quality. Availability: Source code is available at the authors website: http://homepages.cs.ncl.ac.uk/m.j.bell1/annotation. Contact: phillip.lord@newcastle.ac.uk

연구 동기 및 목표

  • 외부 메타데이터나 온톨로지에 의존하지 않고 생물학적 애너테이션 품질을 평가하기 위한 일반적이고 텍스트 중심의 지표를 개발하는 것.
  • 시간이 지남에 따라 단어 빈도 분포의 변화가 애너테이션 품질과 쿠레이션 관행의 변화를 반영하는지 조사하는 것.
  • 언어적 패턴과 지프의 최소 노력 원리를 활용하여 수작업(Swiss-Prot)과 자동(TrEMBL) 애너테이션 품질을 비교하는 것.
  • 단어 빈도의 힘의 법칙 피팅에서 유도된 매개변수 α가 애너테이션 품질의 신뢰할 수 있는 대체 지표로 기능할 수 있는지 평가하는 것.
  • 이 방법이 애너테이션 내에서 저품질 또는 비생물학적 콘텐츠를 탐지하는 데 잠재력을 지니고 있는지 탐색하는 것.

제안 방법

  • 저자들은 1998년부터 2012년까지의 유니프로트KB 버전에서 수집된 모든 자유 텍스트 애너테이션을 추출하였으며, 주로 Swiss-Prot와 TrEMBL에 집중하였다.
  • 모든 애너테이션의 단어 빈도 수를 계산하고, 순위화된 단어 빈도에 대해 힘의 법칙 분포를 피팅하여 척도 매개변수 α를 추정하였다.
  • α 값은 지프의 최소 노력 원리를 통해 해석되었으며, 높은 α 값은 독자가 더 쉽게 이해할 수 있는 예측 가능하고 독자 우환 언어를 의미한다.
  • 저자들은 시간 경과에 따른 α 값의 변화를 비교하여 애너테이션 품질의 추세를 파악하고, 수작업 및 자동 애너테이션 간의 차이를 분석하였다.
  • 단백질 항목의 코hort를 분석하여 초기 애너테이션에서 성숙한 애너테이션으로의 변화 과정에서 α 값이 어떻게 변화하는지 평가하였다.
  • 이 방법은 텍스트 콘텐츠에만 의존하므로, 온톨로지나 증거 코드 사용 여부와 관계없이 자유 텍스트 애너테이션이 있는 모든 데이터베이스에 적용 가능하다.
Figure 1: Outline view of the data extraction process. (1) Initially we download a complete dataset for a given database version in flat file format. (2) We then extract the comment lines (lines beginning with ‘CC’, the comment indicator). (3) We remove comment blocks and properties (as defined in t
Figure 1: Outline view of the data extraction process. (1) Initially we download a complete dataset for a given database version in flat file format. (2) We then extract the comment lines (lines beginning with ‘CC’, the comment indicator). (3) We remove comment blocks and properties (as defined in t

실험 결과

연구 질문

  • RQ1생물학적 애너테이션의 단어 빈도에 대한 힘의 법칙 피팅이 애너테이션 품질의 신뢰할 수 있는 대체 지표로 기능할 수 있는가?
  • RQ2수작업 쿠레이션과 자동 쿠레이션 애너테이션에서 시간이 지남에 따라 단어 빈도 분포에서 유도된 α 매개변수는 어떻게 변화하는가?
  • RQ3α의 변화가 애너테이터의 노력 증가, 독자의 노력 증가와 같은 애너테이션 관행의 변화를 어느 정도 반영하는가?
  • RQ4이 방법은 성숙한 항목이나 새로 추가된 항목의 품질 저하를 탐지할 수 있는가?
  • RQ5α 매개변수는 애너테이션 내에서 비생물학적 또는 저신호 콘텐츠를 식별하는 데 유용한가?

주요 결과

  • Swiss-Prot와 TrEMBL 양쪽에서 α 매개변수가 시간이 지남에 따라 감소하여 독자가 더 많은 노력을 기울여야 하며, 애너테이션 품질이 떨어지고 있음을 시사한다.
  • Swiss-Prot는 초반에 더 높은 α 값을 보였으며, 이는 높은 품질의 독자 친화적 애너테이션을 의미하지만, 데이터 양 증가와 쿠레이션 압박 증가로 인해 시간이 지남에 따라 품질이 악화되었다.
  • TrEMBL는 Swiss-Prot보다 일관되게 낮은 α 값을 보였으며, 이는 자동 애너테이션이 독자 이해보다는 애너테이터의 편의성 향상에 더 최적화되어 있음을 시사한다.
  • 유니프로트KB의 성숙한 항목은 시간이 지남에 따라 α가 서서히 감소하는 경향을 보였으며, 이는 데이터베이스의 성장에 따라 심지어 확립된 애너테이션도 품질이 떨어짐을 의미한다.
  • 최근에 추가된 항목들 역시 시간이 지남에 따라 α 값이 감소하는 경향을 보였으며, 이는 신규 기록의 경우에도 애너테이션 품질이 향상되지 않는다는 것을 시사한다.
  • 이 방법은 저작권 문구와 같은 비생물학적 콘텐츠를 주석 라인 내에서 성공적으로 탐지하여, 이의 오류 탐지 잠재력이 있음을 입증하였다.
Figure 2: Cumulative distributions of words for various Swiss-Prot and TrEMBL versions, shown with logarithmic scales. The size (number of words) is shown along the $X$ axis while the probability is shown on the $Y$ axis. A point on the graph represents the probability that a word will occur $x$ or
Figure 2: Cumulative distributions of words for various Swiss-Prot and TrEMBL versions, shown with logarithmic scales. The size (number of words) is shown along the $X$ axis while the probability is shown on the $Y$ axis. A point on the graph represents the probability that a word will occur $x$ or

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.