[论文解读] Rethinking the production and publication of machine-reusable expressions of research findings
本文介紹了「reborn」——一種預出版框架,通過開放研究知識圖譜(ORKG)將機器可重用的科學知識直接嵌入數據分析工作流程中。與傳統的手動或半自動化後出版提取方法相比,該方法在知識生產過程中即將研究成果結構化為符合FAIR標準、具語義註解的數據,而非在出版後再進行提取,因此實現了更高的準確性、更豐富的知識表達,以及更簡化的技術整合。
Literature is the primary expression of scientific knowledge and an important source of research data. However, scientific knowledge expressed in narrative text documents is not inherently machine reusable. To facilitate knowledge reuse, e.g. for synthesis research, scientific knowledge must be extracted from articles and organized into databases post-publication. The high time costs and inaccuracies associated with completing these activities manually has driven the development of techniques that automate knowledge extraction. Tackling the problem with a different mindset, we propose a pre-publication approach, known as reborn, that ensures scientific knowledge is born reusable, i.e. produced in a machine-reusable format during knowledge production. We implement the approach using the Open Research Knowledge Graph infrastructure for FAIR scientific knowledge organization. We test the approach with three use cases, and discuss the role of publishers and editors in scaling the approach. Our results suggest that the proposed approach is superior compared to classical manual and semi-automated post-publication extraction techniques in terms of knowledge richness and accuracy as well as technological simplicity.
研究动机与目标
- 解決後出版知識提取方法的局限性,如耗時、易出錯且經常不完整。
- 讓研究人員能在數據分析過程中、而非稿件發表後,即在知識創建的時刻產出機器可重用的科學知識。
- 透過使用 ORKG 基礎設施將結構化、語義豐富的數據直接嵌入研究工作流程,提升科學知識的 FAIR 性。
- 減少對手動或半自動化後出版提取的依賴,將知識結構化整合至主要研究流程中。
- 探討出版商與編輯在推動預出版機器可重用知識生產規模化應用中的角色。
提出的方法
- 利用 ORKG 的 Python 與 R 庫,將機器可重用知識生產整合至統計計算環境(如 Python、R)中,實現與資料框的無縫整合。
- 使用 LATEX 為基礎的撰寫方式,搭配 ORKG 模板,在稿件準備期間對研究成果(如資料集、指標、分數)進行語義豐富、機器可讀的註解與結構化。
- 將補充資料(如程式碼、圖表、表格)以結構化 JSON-LD 檔案形式發布,透過 DOI 標記中的「IsSupplementTo」與「HasPart」關係與文章連結。
- 使用 TIB Leibniz 數據管理器作為集中式、持久化的補充資料倉儲,確保長期可及性與互聯結性。
- 透過 ORKG 的 REST 與 SPARQL API,使用文章 DOI 或目錄路徑進行結構化資料的採集,支援生產與沙盒環境。
- 使用 Crossref 標準元數據實現文章與資料之間的雙向互聯結,並可選加入「is-supplemented-by」關係以提升機器可發現性。

实验结果
研究问题
- RQ1與後出版方法相比,預出版階段以機器可讀格式結構化科學知識,是否能提升知識提取的準確性與豐富度?
- RQ2如何在不打擾既有研究實務的情況下,將機器可重用的科學知識直接嵌入數據分析工作流程?
- RQ3出版商與編輯在推動跨學術傳播中預出版知識結構化的規模化採用方面,可扮演何種角色?
- RQ4預出版知識結構化在多大程度上能減少系統性文獻回顧等綜合研究所需耗費的時間與精力?
- RQ5如何利用現有的 DOI 與元數據基礎設施,可靠地發布與發現持久、互操作且符合 FAIR 標準的資料?
主要发现
- reborn 方法在三項真實世界用例中均顯示,其產出的知識比後出版提取更具準確性與更豐富的結構特徵。
- 在分析過程中產出的機器可重用資料,與 ORKG 的 Python 與 R 庫原生相容,可直接載入資料框以供後續分析使用。
- 透過 DOI 連結文章與補充 JSON-LD 資料,可實現可靠、程式化的方式,利用標準 API(如 DataCite 的 REST 接口)發現可採集的研究資料。
- 存放在 TIB Leibniz 數據管理器中的補充資料,依 CC0 1.0 Universal 授權發布,確保開放、持久且可重用的存取。
- 該方法支援 ORKG 資料貢獻在生產與沙盒環境中的部署,但為確保環境間的模板相容性,需將模板與 ORKG 內部識別碼系統分離。
- ORKG 及其組件以 MIT 授權釋出,原始碼與資料可在 GitLab 上公開取得,並透過多個公開 API 提供存取。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。