[Paper Review] A Survey of Large Language Models for Healthcare: from Data, Technology, and Applications to Accountability and Ethics
This survey analyzes how large language models (LLMs) are developed and applied in healthcare, compares them with traditional PLMs, and discusses ethics and evaluation.
The utilization of large language models (LLMs) in the Healthcare domain has generated both excitement and concern due to their ability to effectively respond to freetext queries with certain professional knowledge. This survey outlines the capabilities of the currently developed LLMs for Healthcare and explicates their development process, with the aim of providing an overview of the development roadmap from traditional Pretrained Language Models (PLMs) to LLMs. Specifically, we first explore the potential of LLMs to enhance the efficiency and effectiveness of various Healthcare applications highlighting both the strengths and limitations. Secondly, we conduct a comparison between the previous PLMs and the latest LLMs, as well as comparing various LLMs with each other. Then we summarize related Healthcare training data, training methods, optimization strategies, and usage. Finally, the unique concerns associated with deploying LLMs in Healthcare settings are investigated, particularly regarding fairness, accountability, transparency and ethics. Our survey provide a comprehensive investigation from perspectives of both computer science and Healthcare specialty. Besides the discussion about Healthcare concerns, we supports the computer science community by compiling a collection of open source resources, such as accessible datasets, the latest methodologies, code implementations, and evaluation benchmarks in the Github. Summarily, we contend that a significant paradigm shift is underway, transitioning from PLMs to LLMs. This shift encompasses a move from discriminative AI approaches to generative AI approaches, as well as a shift from model-centered methodologies to data-centered methodologies. Also, we determine that the biggest obstacle of using LLMs in Healthcare are fairness, accountability, transparency and ethics.
Motivation & Objective
- Summarize the development roadmap from pretrained language models (PLMs) to large language models (LLMs) in healthcare.
- Compare PLMs and LLMs and analyze their strengths, limitations, and applications in medical domains.
- Catalog healthcare training data, training methods, optimization strategies, and usage guidelines for LLMs.
- Examine fairness, accountability, transparency, and ethics concerns in deploying healthcare LLMs.
- Provide open-source resources and practical guidance for building private healthcare LLMs.
Proposed method
- Review and synthesize key developments from PLMs to LLMs in healthcare.
- Summarize capabilities and limitations of LLMs across healthcare tasks such as NER, RE, TC, STS, QA, and dialogue.
- Outline data sources, training approaches, optimization strategies, and evaluation methods for healthcare LLMs.
- Discuss fairness, accountability, transparency, and ethics in healthcare LLM deployment.
- Compile open-source datasets, methodologies, code, and benchmarks relevant to healthcare LLMs.
![Figure 1: The development from PLMs to LLMs. GPT-3 [ 17 ] marks a significant milestone in the transition from PLMs to LLMs, signaling the beginning of a new era.](https://ar5iv.labs.arxiv.org/html/2310.05694/assets/Fig1.png)
Experimental results
Research questions
- RQ1What are the capabilities and limitations of current LLMs in healthcare applications?
- RQ2How do PLMs differ from LLMs in healthcare development and usage, and what are the implications for practice?
- RQ3What data, training methods, and evaluation strategies are used for healthcare LLMs, and how do they impact performance and safety?
- RQ4What ethical, fairness, accountability, and transparency concerns arise with healthcare LLMs, and how can they be addressed?
Key findings
- LLMs enable advancements in diverse healthcare tasks including NER, RE, TC, STS, QA, and dialogue generation.
- Med-PaLM 2 achieves high performance on USMLE-style questions, illustrating expert-level potential in medical domains.
- There is a paradigm shift from discriminative PLMs to generative LLMs and from model-centered to data-centered development in healthcare.
- Healthcare LLMs increasingly rely on multimodal data and knowledge graphs to support complex clinical reasoning and reporting.
- The survey provides a collection of open-source datasets, methodologies, code, and benchmarks to support private healthcare LLM development.
- Ethical considerations such as robustness, bias, fairness, accountability, and transparency are analyzed with guidance for responsible deployment.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.