Skip to main content
QUICK REVIEW

[论文解读] A Scalable Data Science Platform for Healthcare and Precision Medicine Research

Jacob McPadden, Thomas J S Durant|arXiv (Cornell University)|Aug 14, 2018
Scientific Computing and Data Management被引用 6
一句话总结

本文提出了一种基于 Hadoop、Apache Storm 和 NiFi 的可扩展、开源的数据科学平台,用于将实时临床数据(如电子健康记录、实验室结果和患者监测数据)整合到统一的数据湖中。该平台实现了精准医疗和计算健康医疗的近实时分析,展示了在大型学术医疗系统中的可行性与性能。

ABSTRACT

Objective: To (1) demonstrate the implementation of a data science platform built on open-source technology within a large, academic healthcare system and (2) describe two computational healthcare applications built on such a platform. Materials and Methods: A data science platform based on several open source technologies was deployed to support real-time, big data workloads. Data acquisition workflows for Apache Storm and NiFi were developed in Java and Python to capture patient monitoring and laboratory data for downstream analytics. Results: The use of emerging data management approaches along with open-source technologies such as Hadoop can be used to create integrated data lakes to store large, real-time data sets. This infrastructure also provides a robust analytics platform where healthcare and biomedical research data can be analyzed in near real-time for precision medicine and computational healthcare use cases. Discussion: The implementation and use of integrated data science platforms offer organizations the opportunity to combine traditional data sets, including data from the electronic health record, with emerging big data sources, such as continuous patient monitoring and real-time laboratory results. These platforms can enable cost-effective and scalable analytics for the information that will be key to the delivery of precision medicine initiatives. Conclusion: Organizations that can take advantage of the technical advances found in data science platforms will have the opportunity to provide comprehensive access to healthcare data for computational healthcare and precision medicine research.

研究动机与目标

  • 在大型学术医疗系统中,利用开源技术设计并部署一个可扩展的数据科学平台。
  • 整合多种数据源,包括电子健康记录、实时实验室结果和持续的患者监测数据。
  • 为精准医疗和计算健康医疗应用实现近实时分析。
  • 展示该平台处理大规模数据工作负载并支持复杂研究工作负载的能力。

提出的方法

  • 基于开源技术(包括 Hadoop 用于分布式存储和处理)部署数据科学平台。
  • 使用 Apache Storm 和 Apache NiFi(用 Java 和 Python 开发)实现数据采集管道,以摄取流式临床数据。
  • 构建统一的数据湖,用于存储异构数据源,包括结构化 EHR 数据和非结构化或半结构化监测数据。
  • 通过事件驱动架构和流处理框架,实现近实时数据处理与分析。
  • 将多个数据源整合到统一的基础设施中,以支持后续研究和临床分析。
  • 利用开放标准和模块化组件,确保在临床和研究系统之间具备可扩展性和互操作性。

实验结果

研究问题

  • RQ1如何在大型学术医疗系统中有效部署一个可扩展、开源的数据平台,以支持实时数据分析?
  • RQ2为实现精准医疗研究,整合实时临床数据流与传统 EHR 数据所需的关键架构组件是什么?
  • RQ3Hadoop、Storm 和 NiFi 等开源技术能否有效管理并处理大规模、异构的临床数据?
  • RQ4此类平台在真实临床工作负载下表现出怎样的性能和可靠性特征?
  • RQ5该平台如何支持精准医疗应用的近实时分析?

主要发现

  • 平台成功使用 Apache Storm 和 NiFi 摄取并处理了来自患者监测和实验室系统的实时数据流。
  • 建立了统一的数据湖,实现了对多样化临床数据源的集中化存储和访问。
  • 该系统在处理学术医疗环境典型的大型实时数据工作负载方面表现出可扩展性和高性能。
  • 将 EHR 数据与流式临床数据集成,实现了精准医疗应用的近实时分析。
  • 使用开源技术实现了成本效益高、可扩展且互操作的数据基础设施,能够支持复杂的研究工作负载。
  • 通过统一异构数据源并支持可扩展分析,该平台为计算健康医疗和精准医疗研究提供了坚实基础。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。