Skip to main content
QUICK REVIEW

[论文解读] The Research Object Suite of Ontologies: Sharing and Exchanging Research Data and Methods on the Open Web

Khalid Belhajjame, Jun Zhao|arXiv (Cornell University)|Jan 17, 2014
Scientific Computing and Data Management参考文献 5被引用 15
一句话总结

本文提出了研究对象(Research Object, RO)系列本体——一种基于语义网的框架,可将研究数据、方法和溯源信息封装为可重用、可共享的容器。通过利用现有词汇表并集成 myExperiment 和 RODL 等工具,RO 模型实现了科学调查的结构化、机器可读打包,显著提升了研究成果的可重现性、可解释性以及长期保存能力。

ABSTRACT

Research in life sciences is increasingly being conducted in a digital and online environment. In particular, life scientists have been pioneers in embracing new computational tools to conduct their investigations. To support the sharing of digital objects produced during such research investigations, we have witnessed in the last few years the emergence of specialized repositories, e.g., DataVerse and FigShare. Such repositories provide users with the means to share and publish datasets that were used or generated in research investigations. While these repositories have proven their usefulness, interpreting and reusing evidence for most research results is a challenging task. Additional contextual descriptions are needed to understand how those results were generated and/or the circumstances under which they were concluded. Because of this, scientists are calling for models that go beyond the publication of datasets to systematically capture the life cycle of scientific investigations and provide a single entry point to access the information about the hypothesis investigated, the datasets used, the experiments carried out, the results of the experiments, the people involved in the research, etc. In this paper we present the Research Object (RO) suite of ontologies, which provide a structured container to encapsulate research data and methods along with essential metadata descriptions. Research Objects are portable units that enable the sharing, preservation, interpretation and reuse of research investigation results. The ontologies we present have been designed in the light of requirements that we gathered from life scientists. They have been built upon existing popular vocabularies to facilitate interoperability. Furthermore, we have developed tools to support the creation and sharing of Research Objects, thereby promoting and facilitating their adoption.

研究动机与目标

  • 解决学术传播中的空白:已发表的文章缺乏足够的上下文信息,难以实现科学结果的可重现性和再利用。
  • 提供一种标准化的、机器可处理的方式,将研究数据、方法、工作流、溯源信息和结论整合为单一可移植的单元。
  • 通过捕获研究调查的完整生命周期(包括假设、实验和结论),支持科学调查的长期保存与可解释性。
  • 通过与既定词汇表对齐,实现与现有存储库和工具(如 DataVerse、FigShare、myExperiment 和 RODL)的互操作性。
  • 通过工具开发和社区参与推动采纳,包括成立 W3C 社区组,以支持持续开发和反馈。

提出的方法

  • 设计核心研究对象本体,用于建模轻量级容器,以聚合研究资源及其元数据。
  • 创建扩展本体,用于科学工作流(wfdesc)、溯源(prov)和时间演化(ro-cd),以捕获执行细节和血缘关系。
  • 将 RO 模型与现有标准(如 W3C PROV、W3C Web Annotation 和 W3C Web Services)集成,确保互操作性。
  • 在 myExperiment 和 RODL 等实际工具中实现 RO 模型,使用户能够创建、管理并共享包含假设、工作流执行和结论元数据的研究对象。
  • 使用 RODL 作为后端存储库,用于存储和管理研究对象,并通过转换服务将工作流定义转换为标准化的 RO 兼容格式。
  • 通过与 BioVel、Scape、Timbus 和 GigaScience 等项目的合作,在真实用例中应用该模型,以验证并优化本体设计。

实验结果

研究问题

  • RQ1如何对科学调查进行封装,以实现完全可重现性和长期保存?
  • RQ2需要何种本体模型,才能语义化地描述研究调查的完整生命周期,包括假设、数据、工作流和结论?
  • RQ3如何扩展现有的研究数据和方法存储库,以支持结构化、机器可读的研究对象共享?
  • RQ4需要哪些技术和组织机制,才能在科学社区中广泛推广研究对象?
  • RQ5如何对研究对象的溯源和版本控制进行建模,以支持可追溯性和审计能力?

主要发现

  • 研究对象本体套件成功实现了可移植、语义丰富的容器,将研究数据、方法和元数据整合为单一可共享单元。
  • 与 myExperiment 和 RODL 的集成使科学家能够显式建模假设、工作流执行和结论,从而创建和管理研究对象。
  • 使用 PROV 和 wfdesc 等既定词汇表,确保了在不同科学工具和存储库之间实现互操作性和可重用性。
  • 该模型已被 BioVel、Scape、Timbus 和 GigaScience 等多个研究社区采纳并扩展,证明了其实际效用和可扩展性。
  • W3C 研究对象社区组的成立已吸引超过 80 名贡献者,表明社区对该标准的浓厚兴趣和强劲发展势头。
  • 该框架通过显式呈现和提供上下文信息(如实验设计和数据血缘),显著提升了科学结果的可重现性和可解释性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。