Skip to main content
QUICK REVIEW

[論文レビュー] Rethinking the production and publication of machine-reusable expressions of research findings

Markus Stocker, Lauren E. Snyder|arXiv (Cornell University)|May 21, 2024
Scientific Computing and Data ManagementDecision Sciences被引用数 3
ひとこと要約

本論文は、オープン・リサーチ・ノリッジ・グラフ(ORKG)を用いて、データ分析ワークフローに直接機械再利用可能な科学的知識を埋め込む『reborn』と呼ばれる、出版前フレームワークを紹介する。研究結果を知識生成段階でFAIRで、意味的にアノテートされたデータとして構造化することにより、出版後の抽出に代わる方法よりも、より高い正確性、より豊富な知識表現、より簡単な技術統合を実現する。

ABSTRACT

Literature is the primary expression of scientific knowledge and an important source of research data. However, scientific knowledge expressed in narrative text documents is not inherently machine reusable. To facilitate knowledge reuse, e.g. for synthesis research, scientific knowledge must be extracted from articles and organized into databases post-publication. The high time costs and inaccuracies associated with completing these activities manually has driven the development of techniques that automate knowledge extraction. Tackling the problem with a different mindset, we propose a pre-publication approach, known as reborn, that ensures scientific knowledge is born reusable, i.e. produced in a machine-reusable format during knowledge production. We implement the approach using the Open Research Knowledge Graph infrastructure for FAIR scientific knowledge organization. We test the approach with three use cases, and discuss the role of publishers and editors in scaling the approach. Our results suggest that the proposed approach is superior compared to classical manual and semi-automated post-publication extraction techniques in terms of knowledge richness and accuracy as well as technological simplicity.

研究の動機と目的

  • 出版後における知識抽出の限界(時間のかかる、誤りが生じやすい、しばしば不完全)を是正すること。
  • 研究者が論文出版後ではなく、データ分析の段階で知識を生成する際、機械再利用可能な科学的知識を生み出せるようにすること。
  • ORKGインfraストラクチャを用いて、研究ワークフローに直接構造的かつ意味的に豊富なデータを埋め込むことで、科学的知識のFAIR性を向上させること。
  • 知識構造化を主たる研究プロセスに統合することで、手動または半自動の出版後抽出に依存するのを減らすこと。
  • 出版者と編集者が、出版前における機械再利用可能な知識生成のスケーラブルな採用を促進する役割を明らかにすること。

提案手法

  • ORKGのPythonおよびRライブラリを用いて、統計計算環境(例:Python、R)に機械再利用可能な知識生成を統合し、データフレームとのシームレスな統合を実現する。
  • ORKGテンプレートを用いたLATEXベースの作成により、研究結果(例:データセット、指標、スコア)を意味的に豊富で機械読取可能なデータとして、論文作成段階でアノテートおよび構造化する。
  • 補足データ(例:コード、図、表)を、DOIメタデータを介して『IsSupplementTo』および『HasPart』関係でリンクした構造化されたJSON-LDファイルとして公開する。
  • 補足データの長期的可視性と相互リンク性を保証するため、TIB レイブニッツ・データマネージャーを中央集約型で永続的なリポジトリとして導入する。
  • 記事のDOIまたはディレクトリパスを用いて、ORKGのRESTおよびSPARQL APIを介して構造化データを収集可能にし、本番環境およびスナップショット環境の両方をサポートする。
  • Crossrefメタデータ基準を用いて、記事とデータの間で双方向の相互リンクを実現し、機械検索性を高めるためにオプションの「is-supplemented-by」関係を提供する。
Figure 1: Scientific knowledge expressed in articles is produced as machine-reusable data in computing environments during the data analysis phase of the research lifecycle. Machine-reusable scientific knowledge is deposited in a data repository as supplementary data of the article and interlinked w
Figure 1: Scientific knowledge expressed in articles is produced as machine-reusable data in computing environments during the data analysis phase of the research lifecycle. Machine-reusable scientific knowledge is deposited in a data repository as supplementary data of the article and interlinked w

実験結果

リサーチクエスチョン

  • RQ1機械読取可能な形式で科学的知識を出版前段階で構造化することで、出版後抽出手法に比べ、知識抽出の正確性と豊かさが向上するか?
  • RQ2既存の研究実践を損なわず、機械再利用可能な科学的知識をデータ分析ワークフローに直接埋め込む方法は何か?
  • RQ3出版者と編集者が、学術コミュニケーション全体にわたり、出版前における知識構造化の採用をスケーラブルに促進する役割を果たせるか?
  • RQ4出版前における知識構造化は、システマティックレビューのような合成研究に要する時間と労力の削減にどの程度寄与するか?
  • RQ5既存のDOIおよびメタデータインfraストラクチャを用いて、永続的で相互運用可能でFAIR資格を満たすデータを信頼性高く公開・発見できるか?

主な発見

  • 3つの実世界のユースケースを通じて、rebornアプローチは、出版後抽出に比べ、より高い正確性とより豊富な構造的豊かさを持つ知識を生み出すことが実証された。
  • 分析段階で生成された機械再利用可能なデータは、natively ORKGのPythonおよびRライブラリと互換性があり、後続の分析に直接データフレームに取り込むことができる。
  • DOIベースの記事と補足JSON-LDデータの相互リンクにより、DataCiteのRESTインターフェースなどの標準APIを用いて、プログラム的かつ信頼性のある研究データの発見が可能になる。
  • TIB レイブニッツ・データマネージャーに寄託された補足データはCC0 1.0ユニバーサルライセンスの下で公開されており、オープンで、永続的かつ再利用可能なアクセスが保証される。
  • 本アプローチは、ORKGの貢献を本番環境およびスナップショット環境にデプロイ可能であるが、環境間でのテンプレート互換性を確保するには、テンプレートをORKGの内部識別子システムから分離する必要がある。
  • ORKGおよびそのコンponentsはMITライセンスの下でリリースされており、ソースコードとデータはGitLabで公開され、複数の公開APIを介して利用可能である。
Figure 2: Display of the research finding published by Gentsch et al. in their Figure 1 as a research contribution in ORKG. The overlay expands on the interlinked R script snippet used to implement the respective data analysis. For an interactive experience, we refer readers to the version published
Figure 2: Display of the research finding published by Gentsch et al. in their Figure 1 as a research contribution in ORKG. The overlay expands on the interlinked R script snippet used to implement the respective data analysis. For an interactive experience, we refer readers to the version published

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。