Skip to main content
QUICK REVIEW

[논문 리뷰] PhyloPythiaS+: A self-training method for the rapid reconstruction of low-ranking taxonomic bins from metagenomes

Ivan Gregor, Johannes Dröge|arXiv (Cornell University)|2014. 06. 27.
Genomics and Phylogenetic Studies참고 문헌 51인용 수 17
한 줄 요약

PhyloPythiaS+는 메타게놈에 대한 조성 기반 분류기의 생성을 자동화하는 자기학습 방법을 도입하여 수동 전문가 코딩을 대체한다. 이는 100배 빠른 k-mer 수세기, 총 실행 시간 3배 감소를 가능하게 하며, 저비용 하드웨어를 사용하여 Gb 크기의 메타게놈으로부터 종 및 분류군 수준의 분류군을 완전 자동으로 고정밀도로 재구성할 수 있도록 한다.

ABSTRACT

Metagenomics is an approach for characterizing environmental microbial communities in situ, it allows their functional and taxonomic characterization and to recover sequences from uncultured taxa. For communities of up to medium diversity, e.g. excluding environments such as soil, this is often achieved by a combination of sequence assembly and binning, where sequences are grouped into 'bins' representing taxa of the underlying microbial community from which they originate. Assignment to low-ranking taxonomic bins is an important challenge for binning methods as is scalability to Gb-sized datasets generated with deep sequencing techniques. One of the best available methods for the recovery of species bins from an individual metagenome sample is the expert-trained PhyloPythiaS package, where a human expert decides on the taxa to incorporate in a composition-based taxonomic metagenome classifier and identifies the 'training' sequences using marker genes directly from the sample. Due to the manual effort involved, this approach does not scale to multiple metagenome samples and requires substantial expertise, which researchers who are new to the area may not have. With these challenges in mind, we have developed PhyloPythiaS+, a successor to our previously described method PhyloPythia(S). The newly developed + component performs the work previously done by the human expert. PhyloPythiaS+ also includes a new k-mer counting algorithm, which accelerated k-mer counting 100-fold and reduced the overall execution time of the software by a factor of three. Our software allows to analyze Gb-sized metagenomes with inexpensive hardware, and to recover species or genera-level bins with low error rates in a fully automated fashion.

연구 동기 및 목표

  • 메타게놈의 분류군 분류에 있어 수동 전문가 코딩의 필요성을 제거하기 위해.
  • 대규모(Gb 크기의) 메타게놈 데이터셋에 대한 확장 가능한 자동 분석을 가능하게 하기 위해.
  • 특수 전문 지식 없이도 고정밀도의 종 및 분류군 수준의 분류군 분리가 가능하도록 하기 위해.
  • 외부 기준 데이터베이스에 대한 의존도를 줄이기 위해 메타게놈 데이터 자체에서 학습하는 자기학습 프레임워크를 개발하기 위해.
  • 마이크로바이옴 연구에서 분류군 분리에 필요한 계산 시간과 자원 요구량을 크게 줄이기 위해.

제안 방법

  • 이 방법은 메타게놈 샘플에서 마커 유전자를 자동으로 식별하여 분류군 분류를 정의하는 자기학습 파이프라인을 사용한다.
  • 데이터 기반 접근을 통해 분류학적으로 유의미한 서열을 식별함으로써 전문가가 훈련 서열을 선택하는 역할을 대체한다.
  • 새로운 k-mer 수세기 알고리즘이 k-mer 빈도 계산을 100배 빠르게 하여 전체 실행 시간을 크게 단축시킨다.
  • 소프트웨어는 k-mer 빈도를 활용하여 구성 기반 분류를 수행하여 서열을 분류군 분류에 할당한다.
  • 표준 하드웨어에서도 효율적으로 실행되도록 설계되어 마이크로바이옴 연구에서의 일상적 사용이 가능하도록 한다.
  • 원시 시퀀싱 데이터에서 분류군 분류에 이르기까지 완전 자동화된 파이프라인을 제공하여 사용자 간섭을 최소화한다.

실험 결과

연구 질문

  • RQ1자기학습 방법이 메타게놈의 분류군 분류기 구축에 있어 수동 전문가 코딩을 대체할 수 있는가?
  • RQ2분류군 분리 정확도를 손상시키지 않고 k-mer 수세기를 얼마나 빠르게 개선할 수 있는가?
  • RQ3자동 분류가 대규모 메타게놈 데이터셋에서 낮은 오류율로 종 및 분류군 수준의 해상도를 달성할 수 있는가?
  • RQ4표준 컴퓨팅 하드웨어만으로도 고정밀도의 낮은 분류군 수준의 분류군 분리를 수행하는 것이 가능한가?
  • RQ5자기학습 분류기의 성능은 전문가가 코딩한 방법인 PhyloPythiaS와 비교해 정확도 및 속도 측면에서 어떻게 다른가?

주요 결과

  • PhyloPythiaS+의 자기학습 접근법은 수동 전문가 코딩을 성공적으로 대체하여 분류 파이프라인의 완전 자동화를 가능하게 하였다.
  • 새로운 k-mer 수세기 알고리즘이 k-mer 빈도 계산을 100배 빠르게 하였다.
  • 소프트웨어의 총 실행 시간은 원래의 PhyloPythiaS 대비 3배 감소하였다.
  • 이 방법은 저비용 하드웨어에서 Gb 크기의 메타게놈 분석이 가능하여 확장성에 크게 기여하였다.
  • 낮은 분류군 수준(종 및 분류군 수준)의 분류군이 낮은 오류율로 복원되었으며, 자동 분류의 높은 정확도를 입증하였다.
  • 특수 전문 지식 없이도 전문가 코딩된 방법과 유사한 높은 성능과 정확도를 유지하면서도, 전문가 코딩의 필요성을 제거하였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.