[论文解读] Text data mining and data quality management for research information systems in the context of open data and open science
本文提出了一种框架,将文本数据挖掘与数据质量管理相结合,以提升开放科学和开放数据环境中的科研信息系统(RIS)性能。通过应用自然语言处理(NLP)技术从非结构化文本中提取实体和关键词,并利用语义技术统一异构数据源,该方法提升了元数据质量,实现了更优的数据集成,并支持科学机构和图书馆中更有效的搜索与发现。
In the implementation and use of research information systems (RIS) in scientific institutions, text data mining and semantic technologies are a key technology for the meaningful use of large amounts of data. It is not the collection of data that is difficult, but the further processing and integration of the data in RIS. Data is usually not uniformly formatted and structured, such as texts and tables that cannot be linked. These include various source systems with their different data formats such as project and publication databases, CERIF and RCD data model, etc. Internal and external data sources continue to develop. On the one hand, they must be constantly synchronized and the results of the data links checked. On the other hand, the texts must be processed in natural language and certain information extracted. Using text data mining, the quality of the metadata is analyzed and this identifies the entities and general keywords. So that the user is supported in the search for interesting research information. The information age makes it easier to store huge amounts of data and increase the number of documents on the internet, in institutions' intranets, in newswires and blogs is overwhelming. Search engines should help to specifically open up these sources of information and make them usable for administrative and research purposes. Against this backdrop, the aim of this paper is to provide an overview of text data mining techniques and the management of successful data quality for RIS in the context of open data and open science in scientific institutions and libraries, as well as to provide ideas for their application. In particular, solutions for the RIS will be presented.
研究动机与目标
- 解决在多个机构系统中整合异构、非结构化且格式不良的研究数据的挑战。
- 通过从非结构化文本中使用文本数据挖掘提取有意义的实体和关键词,提升科研信息系统(RIS)中的数据质量。
- 通过提升不同数据源之间的互操作性与可发现性,支持开放科学和开放数据倡议。
- 为同步和验证RIS中的内部与外部数据源(包括项目与出版物数据库)提供实用解决方案。
- 通过从自然语言文本中提取语义上有意义的信息,丰富元数据,从而增强用户搜索能力。
提出的方法
- 应用自然语言处理(NLP)技术,从研究文档中的非结构化文本中提取实体(例如研究人员、机构、项目)和关键词。
- 使用语义技术将来自不同来源(如CERIF、RCD和机构数据库)的数据映射并对齐至统一的数据模型。
- 对提取的元数据实施数据质量检查,以确保跨系统的一致性、完整性和准确性。
- 将文本挖掘流程集成至RIS中,持续处理并丰富新出版物和项目报告的元数据。
- 利用标准化数据模型和本体,实现机构信息系统之间的互操作性,并减少数据冗余。
- 在内部与外部数据源之间建立同步机制,以保持元数据记录的实时性与一致性。
实验结果
研究问题
- RQ1如何有效处理科研出版物和项目报告中的非结构化文本,以提取高质量、可操作的元数据?
- RQ2在将异构数据源整合至单一科研信息系统时,哪些技术可确保数据质量与一致性?
- RQ3文本数据挖掘如何支持开放科学与开放数据环境中科研成果的发现与关联?
- RQ4语义技术在统一不同RIS数据模型(如CERIF与RCD)的数据方面发挥何种作用?
- RQ5如何在动态科研信息系统中实现内部与外部数据源的持续同步与验证?
主要发现
- 文本数据挖掘通过从非结构化文本中提取实体与关键词,显著提升了元数据质量,从而增强了索引与搜索能力。
- 语义技术通过在统一数据模型下对齐数据,有效实现了来自不同来源(包括项目与出版物数据库)的数据集成。
- 所提出的框架支持内部与外部数据源的同步,减少了数据不一致与重复。
- NLP技术的应用通过为元数据添加语义上有意义的内容,增强了科研信息的可发现性。
- 该方法在真实世界的RIS部署中展现出实际适用性,尤其适用于致力于支持开放科学的科研机构与图书馆。
- 将数据质量管理与文本挖掘相结合,可构建更可靠、可重用的科研信息系统。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。