Skip to main content
QUICK REVIEW

[論文レビュー] Assigning function to protein-protein interactions: a weakly supervised BioBERT based approach using PubMed abstracts

Aparna Elangovan, Melissa B. Davis|arXiv (Cornell University)|Aug 19, 2020
Biomedical Text Mining and Ontologies被引用数 7
ひとこと要約

本論文では、PubMedの要約をマイニングすることで、タンパク質-タンパク質相互作用(PPI)の機能的タイプを自動的に割り当てる弱教師付きのBioBERTベースのモデル、PPI-BioBERTを提案する。1,800万件の要約をスキャンした結果、全体の正確度が46%(アセチル化では87%)を達成し、IntActなどのデータベースにおけるアノテート済みPPIを顕著に拡張した。

ABSTRACT

Motivation: Protein-protein interactions (PPI) are critical to the function of proteins in both normal and diseased cells, and many critical protein functions are mediated by interactions.Knowledge of the nature of these interactions is important for the construction of networks to analyse biological data. However, only a small percentage of PPIs captured in protein interaction databases have annotations of function available, e.g. only 4% of PPI are functionally annotated in the IntAct database. Here, we aim to label the function type of PPIs by extracting relationships described in PubMed abstracts. Method: We create a weakly supervised dataset from the IntAct PPI database containing interacting protein pairs with annotated function and associated abstracts from the PubMed database. We apply a state-of-the-art deep learning technique for biomedical natural language processing tasks, BioBERT, to build a model - dubbed PPI-BioBERT - for identifying the function of PPIs. In order to extract high quality PPI functions at large scale, we use an ensemble of PPI-BioBERT models to improve uncertainty estimation and apply an interaction type-specific threshold to counteract the effects of variations in the number of training samples per interaction type. Results: We scan 18 million PubMed abstracts to automatically identify 3253 new typed PPIs, including phosphorylation and acetylation interactions, with an overall precision of 46% (87% for acetylation) based on a human-reviewed sample. This work demonstrates that analysis of biomedical abstracts for PPI function extraction is a feasible approach to substantially increasing the number of interactions annotated with function captured in online databases.

研究の動機と目的

  • IntActに登録されたPPIのうち4%しか機能的アノテーションが付いていないという、PPIの機能的アノテーションにおける深刻な空白を解消すること。
  • PubMedの要約からのテクスト情報を用いて、リン酸化やアセチル化などのPPIの機能的タイプを自動的に割り当てるスケーラブルな手法を開発すること。
  • 弱教師付き学習とアンサンブル学習を活用することで、PPI機能分類におけるデータ不足とクラス不均衡の問題を克服すること。
  • ディープラーニングフレームワークにおいて、相互作用タイプ別に最適化されたしきい値を適用することで、不確実性の推定と分類性能を向上させること。

提案手法

  • IntActデータベースのPPIとそれに対応するPubMedの要約をリンクすることで、弱教師付きのデータセットを構築する。
  • 既知の機能的タイプを持つPPIのキュレート済みデータセット上でBioBERTを微調整し、PPI-BioBERTモデルを構築する。
  • 予測の不確実性推定と耐性を向上させるために、PPI-BioBERTモデルのアンサンブルを用いる。
  • 訓練データの不均衡に起因する性能の低下を緩和するため、相互作用タイプ別に分類しきい値を適用する。
  • アンサンブルモデルを用いて1,800万件のPubMed要約をスキャンし、新規の機能的タイプが付与されたPPIを抽出する。
  • 精度と信頼性を評価するため、代表的なサンプルを人間によるレビューで検証する。

実験結果

リサーチクエスチョン

  • RQ1PubMedの要約を対象とした弱教師付きディープラーニングは、タンパク質-タンパク質相互作用の機能的タイプを効果的に同定・分類できるか?
  • RQ2限定的なトレーニング例しか存在しない低リソースなPPI機能タイプに対し、BioBERTベースのモデルはどの程度一般化可能か?
  • RQ3アンサンブルモデリングとタイプ別しきい値の適用は、多様なPPI機能カテゴリにわたる予測の信頼性と正確度をどのように向上させるか?
  • RQ4大規模な生物医学文献コーパスに適用した場合、自動PPI機能抽出のスケーラビリティと正確度はどの程度か?

主な発見

  • モデルは1,800万件のPubMed要約から3,253件の新規で機能的タイプが付与されたPPIを効果的に抽出し、アノテート済みPPIを顕著に拡張した。
  • 人間によるレビューによるサンプリングに基づく全体の正確度は46%であり、大規模なスケールでの信頼性ある性能を示した。
  • アセチル化相互作用では正確度が87%に達し、明確に表現された特定の機能的タイプに対して優れた性能を示した。
  • アンサンブルモデリングの適用により、不確実性の推定が向上し、信頼度の低い予測におけるモデルの信頼性が向上した。
  • タイプ別しきい値の適用により、相互作用タイプごとのトレーニングサンプルの不均衡に起因する性能のばらつきが効果的に緩和された。
  • 本手法は、PPIの機能的アノテーションを大規模に実施する上で実用的であり、手作業によるキュレーションの代替手段としてのスケーラブルな解決策を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。