[论文解读] OpenExtract: Automated Data Extraction for Systematic Reviews in Health
OpenExtract 是一个开源管道,使用大语言模型自动提取健康领域系统综述数据,实现与人工研究者相比 precision 和 recall 均高于 0.8。
This study presents OpenExtract, an open-source pipeline for automated data extraction in large-scale systematic literature reviews. The pipeline queries large language models (LLMs) to predict data entries based on relevant sections of scientific articles. To test the efficacy of OpenExtract, we apply it to a systematic literature review in digital health and compare its outputs with those of human researchers. OpenExtract achieves precision and recall scores of > 0.8 in this task, indicating that it can be effective at extracting data automatically and efficiently. OpenExtract: https://github.com/JimAchterbergLUMC/OpenExtract.
研究动机与目标
- 健全健康系统综述中可扩展数据提取的需求动机。
- 提出一个使用 LLMs 自动化数据提取的开源管道。
- 在数字健康系统综述上将 OpenExtract 与人工研究者进行对比评估。
提出的方法
- OpenExtract 向大型语言模型查询,以从相关论文部分预测数据条目。
- 该管道将通常由研究者执行的数据提取任务自动化。
- 与人工研究者在数字健康系统综述上的对比评估。
实验结果
研究问题
- RQ1一个自动化管道是否能准确提取健康系统综述中预定义的数据条目?
- RQ2在数据提取任务的准确性和召回率方面,OpenExtract 与人类研究者相比如何?
- RQ3开源方法在大规模系统综述中是否有效?
主要发现
- OpenExtract 在提取数字健康系统综述数据条目时实现了 > 0.8 的 precision 和 recall。
- 管道在输出的数据提取结果方面展示出与人类研究者相当的效率。
- 本研究验证 OpenExtract 作为健康领域大规模系统综述的可行自动化工具。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。