Skip to main content
QUICK REVIEW

[論文レビュー] Biolink Model: A Universal Schema for Knowledge Graphs in Clinical, Biomedical, and Translational Science

Deepak Unni, Sierra Moxon|arXiv (Cornell University)|Mar 25, 2022
Biomedical Text Mining and Ontologies参考文献 11被引用数 6
ひとこと要約

Biolink Modelは、生物医学、臨床および翻訳的科学分野における標準化されたオープンソースの知識グラフスキーマを提唱し、エンティティ(例:遺伝子、疾患、化学物質)および関係(述語)の階層的オントロジーを通じて多様なデータソースを統合する。これにより、相互運用性、再利用可能なデータ統合、および知識発見の向上が可能となり、Biomedical Data Translator Consortium や Monarch Initiative などのイニシアチブにおいて、データセット間の推論能力とデータ再利用性が顕著に向上する。

ABSTRACT

Within clinical, biomedical, and translational science, an increasing number of projects are adopting graphs for knowledge representation. Graph-based data models elucidate the interconnectedness between core biomedical concepts, enable data structures to be easily updated, and support intuitive queries, visualizations, and inference algorithms. However, knowledge discovery across these "knowledge graphs" (KGs) has remained difficult. Data set heterogeneity and complexity; the proliferation of ad hoc data formats; poor compliance with guidelines on findability, accessibility, interoperability, and reusability; and, in particular, the lack of a universally-accepted, open-access model for standardization across biomedical KGs has left the task of reconciling data sources to downstream consumers. Biolink Model is an open source data model that can be used to formalize the relationships between data structures in translational science. It incorporates object-oriented classification and graph-oriented features. The core of the model is a set of hierarchical, interconnected classes (or categories) and relationships between them (or predicates), representing biomedical entities such as gene, disease, chemical, anatomical structure, and phenotype. The model provides class and edge attributes and associations that guide how entities should relate to one another. Here, we highlight the need for a standardized data model for KGs, describe Biolink Model, and compare it with other models. We demonstrate the utility of Biolink Model in various initiatives, including the Biomedical Data Translator Consortium and the Monarch Initiative, and show how it has supported easier integration and interoperability of biomedical KGs, bringing together knowledge from multiple sources and helping to realize the goals of translational science.

研究の動機と目的

  • 生物医学および臨床研究における知識グラフのための普遍的でオープンアクセスのデータモデルの欠如に対処すること。
  • 一時的なデータフォーマットやFAIR準拠の不足によるデータの異種性と断片化を低減すること。
  • 複数の生物医学的知識グラフおよびデータソース間でのシームレスな統合と相互運用性を実現すること。
  • コアな生物医学的エンティティ間の関係を形式化することで、翻訳的科学における直感的なクエリ、可視化、推論を支援すること。
  • 多様な研究イニシャチブにわたる知識表現の標準化を可能にする再利用可能で拡張可能なフレームワークを提供すること。

提案手法

  • 遺伝子、疾患、化学物質、解剖的構造などの生物医学的エンティティ(クラス)の階層的でオブジェクト指向のオントロジーを定義すること。
  • エンティティ間の標準化された関係(述語)を確立し、属性や関連付けを含め、意味的モデリングをガイドすること。
  • 複雑で相互接続された生物医学的知識を機械可読形式で表現するためのグラフ指向の機能を統合すること。
  • 複数の知識グラフプロジェクトやデータ統合パイプラインにわたる拡張可能で再利用可能な形でモデルを設計すること。
  • コミュニティの採用と長期的な保守性を確保するため、オープンソースフレームワークとしてモデルを実装すること。
  • Biomedical Data Translator Consortium や Monarch Initiative などの大規模なイニシャチブへの統合を通じて、モデルの妥当性を検証すること。

実験結果

リサーチクエスチョン

  • RQ1普遍的スキーマは、多様な生物医学的知識グラフにおける知識表現をどのように標準化できるか?
  • RQ2共通のデータモデルは、翻訳的科学における相互運用性とデータ統合をどの程度向上させるか?
  • RQ3統一されたオントロジーは、データの異種性を低減し、FAIR(検索可能、アクセス可能、相互運用可能、再利用可能)準拠を向上させることができるか?
  • RQ4Biolink Modelは、実世界の生物医学的応用において、データセット間のクエリおよび推論をどの程度効果的に支援できるか?
  • RQ5標準化されたスキーマは、大規模な生物医学的研究における知識発見とデータ再利用にどのような影響を及ぼすか?

主な発見

  • Biolink Modelは、遺伝子、疾患、化学物質などのコアな生物医学的エンティティ間の関係を標準化する包括的で再利用可能なスキーマを提供する。
  • このモデルは、Biomedical Data Translator Consortium や Monarch Initiative などの複数のソースからの知識統合をシームレスに可能にする。
  • エンティティクラスと関係を形式化することで、データの相互運用性が顕著に向上し、知識グラフ統合の複雑さが低減される。
  • Biolink Modelの採用は、分散型の知識グラフにわたる生物医学的データの再利用可能性と検索可能性を向上させた。
  • 一貫性のある機械処理可能な知識表現を提供することで、このモデルは、意味的クエリや推論を含む高度な分析ワークフローを支援する。
  • このモデルのオープンソース性は、複数の研究イニシャチブにわたるコミュニティの採用と長期的持続可能性を促進した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。