Skip to main content
QUICK REVIEW

[Paper Review] DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models

Jiancheng Ye, Sophie Bronstein|ArXiv.org|Jun 2, 2025
Artificial Intelligence in Healthcare and Education3 citations
TL;DR

A survey of DeepSeek-R1, an open-source LLM for healthcare, detailing its architecture, capabilities, clinical applications, benchmarks, risks, and governance implications.

ABSTRACT

DeepSeek-R1 is a cutting-edge open-source large language model (LLM) developed by DeepSeek, showcasing advanced reasoning capabilities through a hybrid architecture that integrates mixture of experts (MoE), chain of thought (CoT) reasoning, and reinforcement learning. Released under the permissive MIT license, DeepSeek-R1 offers a transparent and cost-effective alternative to proprietary models like GPT-4o and Claude-3 Opus; it excels in structured problem-solving domains such as mathematics, healthcare diagnostics, code generation, and pharmaceutical research. The model demonstrates competitive performance on benchmarks like the United States Medical Licensing Examination (USMLE) and American Invitational Mathematics Examination (AIME), with strong results in pediatric and ophthalmologic clinical decision support tasks. Its architecture enables efficient inference while preserving reasoning depth, making it suitable for deployment in resource-constrained settings. However, DeepSeek-R1 also exhibits increased vulnerability to bias, misinformation, adversarial manipulation, and safety failures - especially in multilingual and ethically sensitive contexts. This survey highlights the model's strengths, including interpretability, scalability, and adaptability, alongside its limitations in general language fluency and safety alignment. Future research priorities include improving bias mitigation, natural language comprehension, domain-specific validation, and regulatory compliance. Overall, DeepSeek-R1 represents a major advance in open, scalable AI, underscoring the need for collaborative governance to ensure responsible and equitable deployment.

Motivation & Objective

  • Assess the capabilities of an open-source LLM (DeepSeek-R1) in healthcare tasks.
  • Characterize the architectural design (Mixture of Experts, CoT, RL) and its implications for cost and inference.
  • Evaluate performance on healthcare and domain benchmarks (e.g., USMLE) and identify clinical decision-support strengths.
  • Identify safety, bias, and misinformation risks, especially in multilingual and ethically sensitive contexts.
  • Outline governance, regulatory, and deployment considerations for open-source medical LLMs.

Proposed method

  • Describe the DeepSeek-R1 hybrid architecture combining mixture of experts, chain-of-thought reasoning, and reinforcement learning.
  • Analyze performance on standard benchmarks (USMLE, AIME) and assess domain-specific capabilities.
  • Evaluate interpretability, scalability, and inference efficiency in resource-constrained settings.
  • Assess safety, bias, misinformation, adversarial risk, and multilingual/ethical challenges.
  • Discuss regulatory compliance and governance considerations for open-source medical LLM deployment.

Experimental results

Research questions

  • RQ1What capabilities does DeepSeek-R1 demonstrate in healthcare tasks and structured problem solving?
  • RQ2What are the strengths and limitations of DeepSeek-R1 in medical domains and decision support?
  • RQ3How does DeepSeek-R1 perform on medical and mathematical benchmarks such as USMLE and AIME?
  • RQ4What safety, bias, misinformation, and adversarial risks affect DeepSeek-R1, particularly in multilingual or ethically sensitive contexts?
  • RQ5What governance, regulatory, and deployment considerations are needed for open-source medical LLMs?

Key findings

  • DeepSeek-R1 demonstrates competitive performance on benchmarks like USMLE and AIME.
  • The model shows strong results in pediatric and ophthalmologic clinical decision support tasks.
  • It offers interpretability, scalability, and efficient inference suitable for resource-constrained settings.
  • There are elevated risks of bias, misinformation, adversarial manipulation, and safety failures, especially in multilingual and sensitive contexts.
  • Open-source licensing (MIT) and hybrid architecture support transparency and cost-effectiveness, underscoring governance needs.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.