Skip to main content
QUICK REVIEW

[논문 리뷰] Biolink Model: A Universal Schema for Knowledge Graphs in Clinical, Biomedical, and Translational Science

Deepak Unni, Sierra Moxon|arXiv (Cornell University)|2022. 03. 25.
Biomedical Text Mining and Ontologies참고 문헌 11인용 수 6
한 줄 요약

Biolink Model는 생물의학, 임상 및 번역 과학 분야에서 표준화되고 오픈소스인 지식 그래프 스키마를 제안하여, 유전자, 질병, 화학물질 등의 엔터티 계층 구조 온톨로지와 관계(예: 술어)를 통해 다양한 데이터 소스를 통합한다. 이는 상호운용성, 재사용 가능한 데이터 통합 및 지식 탐색의 향상을 가능하게 하며, Biomedical Data Translator Consortium와 Monarch Initiative와 같은 주요 연구 이니셔티브에서 심각하게 향상된 데이터셋 간 추론과 데이터 재사용을 이끌어낸다.

ABSTRACT

Within clinical, biomedical, and translational science, an increasing number of projects are adopting graphs for knowledge representation. Graph-based data models elucidate the interconnectedness between core biomedical concepts, enable data structures to be easily updated, and support intuitive queries, visualizations, and inference algorithms. However, knowledge discovery across these "knowledge graphs" (KGs) has remained difficult. Data set heterogeneity and complexity; the proliferation of ad hoc data formats; poor compliance with guidelines on findability, accessibility, interoperability, and reusability; and, in particular, the lack of a universally-accepted, open-access model for standardization across biomedical KGs has left the task of reconciling data sources to downstream consumers. Biolink Model is an open source data model that can be used to formalize the relationships between data structures in translational science. It incorporates object-oriented classification and graph-oriented features. The core of the model is a set of hierarchical, interconnected classes (or categories) and relationships between them (or predicates), representing biomedical entities such as gene, disease, chemical, anatomical structure, and phenotype. The model provides class and edge attributes and associations that guide how entities should relate to one another. Here, we highlight the need for a standardized data model for KGs, describe Biolink Model, and compare it with other models. We demonstrate the utility of Biolink Model in various initiatives, including the Biomedical Data Translator Consortium and the Monarch Initiative, and show how it has supported easier integration and interoperability of biomedical KGs, bringing together knowledge from multiple sources and helping to realize the goals of translational science.

연구 동기 및 목표

  • 생물의학 및 임상 연구 분야에서 지식 그래프를 위한 유니버설이고 오픈 액세스 데이터 모델의 부족을 해결하기 위해.
  • 임의의 데이터 포맷과 낮은 FAIR 준수로 인한 데이터 이질성과 분산 문제를 줄이기 위해.
  • 다양한 생물의학 지식 그래프 및 데이터 소스 간의 원활한 통합과 상호운용성을 보장하기 위해.
  • 핵심 생물의학 엔터티 간의 관계를 체계화함으로써 번역 과학에서 직관적인 질의, 시각화 및 추론을 지원하기 위해.
  • 다양한 연구 이니셔티브에서 지식 표현을 표준화하기 위한 재사용 가능하고 확장 가능한 프레임워크를 제공하기 위해.

제안 방법

  • 유전자, 질병, 화학물질, 해부학적 구조와 같은 생물의학 엔터티(클래스)의 계층적이고 객체 지향적인 온톨로지 정의하기.
  • 엔터티 간의 표준화된 관계(술어) 집합 수립 및 의미 모델링을 안내하기 위해 속성과 연결 고리를 포함하기.
  • 복잡하고 상호 연결된 생물의학 지식을 기계로 읽을 수 있는 형식으로 표현하기 위해 그래프 중심 기능 통합하기.
  • 다양한 지식 그래프 프로젝트와 데이터 통합 파이프라인에서 확장 가능하고 재사용 가능한 방식으로 모델 설계하기.
  • 커뮤니티의 채택과 장기적인 유지보수를 보장하기 위해 오픈소스 프레임워크로 모델 구현하기.
  • Biomedical Data Translator Consortium와 Monarch Initiative와 같은 대규모 연구 이니셔티브에 통합하여 모델 검증하기.

실험 결과

연구 질문

  • RQ1유니버설 스키마는 다양한 생물의학 지식 그래프 간의 지식 표현을 어떻게 표준화할 수 있는가?
  • RQ2공통 데이터 모델은 번역 과학 분야에서 상호운용성과 데이터 통합에 얼마나 효과적인가?
  • RQ3통합 온톨로지가 데이터 이질성을 줄이고 FAIR(발견 가능, 접근 가능, 상호운용 가능, 재사용 가능) 준수를 향상시키는가?
  • RQ4Biolink Model은 실제 생물의학 응용에서 데이터셋 간 질의 및 추론을 얼마나 효과적으로 지원하는가?
  • RQ5표준화된 스키마는 대규모 생물의학 연구에서 지식 탐색과 데이터 재사용에 어떤 영향을 미치는가?

주요 결과

  • Biolink Model은 유전자, 질병, 화학물질와 같은 핵심 생물의학 엔터티 간의 관계를 표준화하는 포괄적이고 재사용 가능한 스키마를 제공한다.
  • 이 모델은 Biomedical Data Translator Consortium와 Monarch Initiative와 같은 다양한 소스의 지식을 원활하게 통합할 수 있도록 한다.
  • 엔터티 클래스와 관계를 체계화함으로써 모델은 데이터 상호운용성을 크게 향상시키고 지식 그래프 통합의 복잡성을 감소시킨다.
  • Biolink Model의 도입은 분산된 지식 그래프 간 생물의학 데이터의 재사용성과 탐색 가능성 향상에 기여했다.
  • 일致하고 기계로 처리 가능한 지식 표현을 제공함으로써 모델은 의미론적 질의 및 추론을 포함한 고급 분석 워크플로우를 지원한다.
  • 모델의 오픈소스 성격은 다양한 연구 이니셔티브에서 커뮤니티의 채택과 장기적 지속 가능성을 촉진했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.