Skip to main content
QUICK REVIEW

[Paper Review] A new paradigm for accelerating clinical data science at Stanford Medicine

Somalee Datta, Jose Posada|arXiv (Cornell University)|Mar 17, 2020
Data Quality and Management53 references33 citations
TL;DR

The paper proposes a secure, standardized, anonymized Big Data platform at Stanford Medicine that integrates with a secure analytics facility to accelerate clinical data science. It emphasizes reproducibility and privacy to enable faster, broader research.

ABSTRACT

Stanford Medicine is building a new data platform for our academic research community to do better clinical data science. Hospitals have a large amount of patient data and researchers have demonstrated the ability to reuse that data and AI approaches to derive novel insights, support patient care, and improve care quality. However, the traditional data warehouse and Honest Broker approaches that are in current use, are not scalable. We are establishing a new secure Big Data platform that aims to reduce time to access and analyze data. In this platform, data is anonymized to preserve patient data privacy and made available preparatory to Institutional Review Board (IRB) submission. Furthermore, the data is standardized such that analysis done at Stanford can be replicated elsewhere using the same analytical code and clinical concepts. Finally, the analytics data warehouse integrates with a secure data science computational facility to support large scale data analytics. The ecosystem is designed to bring the modern data science community to highly sensitive clinical data in a secure and collaborative big data analytics environment with a goal to enable bigger, better and faster science.

Motivation & Objective

  • Motivate the need for scalable clinical data science on Stanford Medicine’s patient data assets.
  • Describe limitations of traditional data warehouses and Honest Broker models in scaling research.
  • Propose a new secure data platform design that speeds access, preserves privacy, and standardizes concepts for reproducibility.
  • Outline how the platform integrates anonymization, IRB preparation, standardization, and secure analytics to enable collaborative big data work.

Proposed method

  • Propose a new data platform architecture combining anonymization, standardization, and secure analytics.
  • Describe processes to prepare data for IRB submission while preserving privacy.
  • Define standard clinical concepts to enable cross-site reproducibility of analyses.
  • Explain integration between an analytics data warehouse and a secure computational facility.
  • Argue for enabling collaboration among the modern data science community on sensitive clinical data.

Experimental results

Research questions

  • RQ1How can Stanford Medicine redesign data infrastructure to reduce time to access and analyze clinical data?
  • RQ2In what ways can anonymization, standardization, and IRB-ready preparation improve reproducibility and safety of clinical data analyses?
  • RQ3What architectural components are needed to securely integrate data storage with large-scale analytics while enabling collaboration?

Key findings

  • A secure Big Data platform is proposed to reduce time to access and analyze data.
  • Data should be anonymized to preserve patient privacy prior to IRB submission.
  • Data should be standardized so analyses can be replicated using the same analytical code and clinical concepts.
  • An analytics data warehouse should integrate with a secure data science computational facility for large-scale analytics.
  • The ecosystem aims to bring modern data science to sensitive clinical data in a secure, collaborative environment to enable larger, better, faster science.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.