Skip to main content
QUICK REVIEW

[論文レビュー] OptimOTU: Taxonomically aware OTU clustering with optimized thresholds and a bioinformatics workflow for metabarcoding data

Brendan Furneaux, Sten Anslan|ArXiv.org|Feb 14, 2025
Environmental DNA in Biodiversity Studies被引用数 3
ひとこと要約

OptimOTU は、低階層ごとの閾値を最適化し、完全な Illumina メタバーコーディングワークフローを統合する分類学的に配慮した OTU クラスタリングアルゴリズムを導入し、 大規模で多様なデータセットの処理を可能にします。

ABSTRACT

To turn environmentally derived metabarcoding data into community matrices for ecological analysis, sequences must first be clustered into operational taxonomic units (OTUs). This task is particularly complex for data including large numbers of taxa with incomplete reference libraries. OptimOTU offers a taxonomically aware approach to OTU clustering. It uses a set of taxonomically identified reference sequences to choose optimal genetic distance thresholds for grouping each ancestor taxon into clusters which most closely match its descendant taxa. Then, query sequences are clustered according to preliminary taxonomic identifications and the optimized thresholds for their ancestor taxon. The process follows the taxonomic hierarchy, resulting in a full taxonomic classification of all the query sequences into named taxonomic groups as well as placeholder "pseudotaxa" which accommodate the sequences that could not be classified to a named taxon at the corresponding rank. The OptimOTU clustering algorithm is implemented as an R package, with computationally intensive steps implemented in C++ for speed, and incorporating open-source libraries for pairwise sequence alignment. Distances may also be calculated externally, and may be read from a UNIX pipe, allowing clustering of large datasets where the full distance matrix would be inconveniently large to store in memory. The OptimOTU bioinformatics pipeline includes a full workflow for paired-end Illumina sequencing data that incorporates quality filtering, denoising, artifact removal, taxonomic classification, and OTU clustering with OptimOTU. The OptimOTU pipeline is developed for use on high performance computing clusters, and scales to datasets with millions of reads per sample, and tens of thousands of samples.

研究の動機と目的

  • Taxon-specific な遺伝的変異と不完全な参照ライブラリを考慮して OTU クラスタリングの改善を動機づける。
  • 祖先分類群ごとにクラスタリング閾値を最適化するアルゴリズムを開発し、分類とより良く一致させる。
  • 生データから分類学的に情報を得た OTU およびプレースホルダ疑似分類群を含む、完全でスケーラブルなパイプラインを提供する。
  • 既存の分類識別ツールおよびオープンソースの距離測度との統合を可能にし、効率性を高める。

提案手法

  • 三つのフェーズを持つ OptimOTU クラスタリングアルゴリズムを導入:閾値最適化、予備分類学的同定、階層的クラスタリング。
  • AMI(修正相互情報量)を用いてランク別のカット閾値を決定するため、ランク間の分類分割を比較して閾値を最適化する。
  • 閉域参照法と新規デノボ法を組み合わせた分類学的ガイド付き階層を用いてクエリをクラスタリングし、命名された分類群と疑似分類群を生成する。
  • 内部の複数の手法(ハミング、Edlib、WFA2)または外部距離行列による距離計算を実装し、複雑なマーカー(例:ITS)では SPEED のために USEARCH 統合を選択可能とする。
  • デフォルトのツリー型クラスタリングアルゴリズムを提供し、同時実行、マージ、階層型モードを含む並列化戦略を採用する。
  • Illumina ペアエンドデータを処理する end-to-end OptimOTU パイプラインにクラスタリングを組み込み、品質フィルタリング、デノイジング、キメラ除去、分類学的ガイド付きクラスタリングを含む。

実験結果

リサーチクエスチョン

  • RQ1分類階層ごとの最適化閾値を用いた分類学的に情報を活用した OTU クラスタリングは、複数の分類群にまたがる単一閾値法よりも精度を向上させるか。
  • RQ2予備的な分類同定を組み込むことで、参照ライブラリが不完全なデータセットのクラスタリング効率と精度にどのような影響を与えるか。
  • RQ3多数のリードと多数のサンプルを含む大規模なメタバーコーディングデータセットにおける OptimOTU の性能とスケーラビリティは、伝統的なワークフローと比較してどうか。
  • RQ4異なる距離計算およびクラスタリング構成が、得られる OTU 分割と下流の生態学的分析にどのような影響を与えるか。

主な発見

  • OptimOTU は祖先分類群ごとに最適化された閾値を用いた分類学的に導かれたクラスタリングを実現し、階層間の分類と整合性を向上させる。
  • パイプラインは品質フィルタリング、デノイジング、アーティファクト除去、階層的クラスタリングを統合し、必要に応じて命名された分類群と疑似分類群を生成する。
  • クラスタリングは大規模データセットをサポートし HPC クラスター上での処理を可能にし、デフォルトはツリー型アルゴリズム、複数の並列化戦略を採用している。
  • 距離は内部計算または外部ソースから取得でき、ITS のような複雑なマーカーで SPEED を高める USEARCH ベースのオプションを含む。
  • このワークフローは raw reads から OTU レベルの分類学的割り当てまでの Illumina ペアエンド全体パイプラインを提供し、真菌 ITS2 および節足動物 COI 分析に適している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。