Skip to main content
QUICK REVIEW

[論文レビュー] Knowledge-based Integration of Multi-Omic Datasets with Anansi: Annotation-based Analysis of Specific Interactions

Thomaz F. S. Bastiaanssen, Thomas P. Quinn|arXiv (Cornell University)|May 18, 2023
Bioinformatics and Genomic Networks被引用数 5
ひとこと要約

この論文では、KEGGなどの知識データベースを用いて対比較的関連を制約することで、多オミクスデータ統合の解釈可能性と統計的パワーを向上させるRパッケージanansiを紹介する。生物学的に妥当な相互作用に限定することで、anansiは誤発見率の補正を軽減し、宿主-微生物系においてより生物学的文脈を踏まえた効率的な有意な代謝物-機能関連を同定する。

ABSTRACT

Motivation: Studies including more than one type of 'omics data sets are becoming more prevalent. Integrating these data sets can be a way to solidify findings and even to make new discoveries. However, integrating multi-omics data sets is challenging. Typically, data sets are integrated by performing an all-vs-all correlation analysis, where each feature of the first data set is correlated to each feature of the second data set. However, all-vs-all association testing produces unstructured results that are hard to interpret, and involves potentially unnecessary hypothesis testing that reduces statistical power due to false discovery rate (FDR) adjustment. Implementation: Here, we present the anansi framework, and accompanying R package, as a way to improve upon all-vs-all association analysis. We take a knowledge-based approach where external databases like KEGG are used to constrain the all-vs-all association hypothesis space, only considering pairwise associations that are a priori known to occur. This produces structured results that are easier to interpret, and increases statistical power by skipping unnecessary hypothesis tests. In this paper, we present the anansi framework and demonstrate its application to learn metabolite-function interactions in the context of host-microbe interactions. We further extend our framework beyond pairwise association testing to differential association testing, and show how anansi can be used to identify associations that differ in strength or degree based on sample covariates such as case/control status. Availability: https://github.com/thomazbastiaanssen/anansi

研究の動機と目的

  • 多オミクス研究における全対比較相関解析から得られる非構造的かつ高次元の結果の解釈の難しさに対処すること。
  • 事前の生物学的知識を活用することで、不要な仮説検定を減らし、統計的パワーを向上させること。
  • 特に宿主-マイクロバイオーム系において、特徴量間の生物学的に意味のある相互作用の同定を可能にすること。
  • 単なる相関を超えて、症例/対照群状態などのサンプル共変数に基づく差分関連解析への応用を可能にすること。
  • 既知の分子相互作用ネットワークに結果を埋め込むことで、検証可能な生物学的仮説を生成するフレームワークを提供すること。

提案手法

  • anansiフレームワークは、外部の知識データベース(例:KEGG、HMDB)を用いて事前に妥当な特徴量ペアを定義し、仮説空間を生物学的に妥当な相互作用に限定する。
  • 異なるオミクス層の特徴量間の既知の生物学的相互作用を符号化した知識ベースの隣接行列を構築する。
  • 全対比較相関検定を、隣接行列に含まれるペアに限定したターゲット検定に置き換え、検定回数とFDR補正の負担を削減する。
  • 標準的な相関検定に加え、健康群と疾患群などのサンプル群間で関連性を比較する差分関連検定もサポートする。
  • 統計的有意性はFDR補正済みp値を用いて評価するが、あくまで事前に生物学的に妥当な相互作用に限って適用する。
  • オープンソースのRパッケージとして実装されており、可視化機能と既存の多オミクスワークフローへの統合をサポートする。

実験結果

リサーチクエスチョン

  • RQ1知識ベースの制約を用いた対比較的関連の制限は、多オミクス統合結果の解釈可能性を向上させるか?
  • RQ2生物学的知識を用いて仮説の数を減らすことで、多オミクス研究における統計的パワーが向上するか?
  • RQ3anansiは疾患状態などの生物学的共変数に応じて、代謝物-機能ペアの差分関連を検出できるか?
  • RQ4標準的な全対比較相関アプローチと比較して、anansiは宿主-マイクロバイオーム系において生物学的に意味のある相互作用をどれほど効果的に同定できるか?
  • RQ5知識データベースの制限が、anansiフレームワークの感度と新規発見可能性にどの程度影響を及ぼすか?

主な発見

  • anansiは、KEGGの既知の代謝経路に沿った構造的で生物学的に解釈可能な代謝物-機能関連を、宿主-マイクロバイオームデータセットで同定した。
  • 既知の相互作用に限定して検定することで、仮説検定の回数を削減し、統計的パワーを向上させるとともにFDR補正の負担を軽減した。
  • フレームワークは、治療群ごとの代謝物-機能ペア間の顕著な不連続な関連を検出し、FDR補正済みp値を用い、色分けされたヒートマップで可視化した。
  • この手法により、特定の微生物遺伝子機能が特定の代謝物と強く相関していることが明らかになったが、これは宿主-代謝物相互作用の文脈において、サンプル共変数に依存した性質を示した。
  • 既存の生物学的ネットワークに結果を統合することで、anansiは解釈可能性を向上させ、研究者が微生物代謝機能に関する検証可能な仮説を立てられるようにした。
  • このアプローチは、非構造的な全対比較解析では見過ごされたり、ぼやけてしまう関連を効果的に同定できることを示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。