Skip to main content
QUICK REVIEW

[Paper Review] Prompt engineering paradigms for medical applications: scoping review and recommendations for better practices

Jamil Zaghir, Marco Naguib|arXiv (Cornell University)|May 2, 2024
Biomedical and Engineering Education7 citations
TL;DR

A scoping review of 114 medical prompt engineering studies (2022–2024) covering PD, PL, and PT, with recommendations to standardize terminology and practices.

ABSTRACT

Prompt engineering is crucial for harnessing the potential of large language models (LLMs), especially in the medical domain where specialized terminology and phrasing is used. However, the efficacy of prompt engineering in the medical domain remains to be explored. In this work, 114 recent studies (2022-2024) applying prompt engineering in medicine, covering prompt learning (PL), prompt tuning (PT), and prompt design (PD) are reviewed. PD is the most prevalent (78 articles). In 12 papers, PD, PL, and PT terms were used interchangeably. ChatGPT is the most commonly used LLM, with seven papers using it for processing sensitive clinical data. Chain-of-Thought emerges as the most common prompt engineering technique. While PL and PT articles typically provide a baseline for evaluating prompt-based approaches, 64% of PD studies lack non-prompt-related baselines. We provide tables and figures summarizing existing work, and reporting recommendations to guide future research contributions.

Motivation & Objective

  • Assess how prompt engineering is applied in medicine across PD, PL, and PT.
  • Identify prevailing paradigms, techniques, and terminologies used in recent medical LLM work.
  • Provide recommendations to improve rigor, baselines, and reporting in future work.

Proposed method

  • Systematic scoping of 114 studies from 2022–2024 applying prompt engineering in medicine.
  • Categorization of studies by prompt design (PD), prompt learning (PL), and prompt tuning (PT).
  • Quantitative synthesis of usage patterns (e.g., prevalence of PD, usage of ChatGPT, common techniques).
  • Extraction of baseline practices and reporting gaps to inform recommendations.

Experimental results

Research questions

  • RQ1What are the dominant prompt engineering paradigms used in medical applications (PD, PL, PT)?
  • RQ2How are these paradigms applied across studies, and what baselines are used for evaluation?
  • RQ3What common techniques (e.g., Chain-of-Thought) are employed, and what LLMs are used?
  • RQ4What gaps exist in reporting and terminology that affect comparability and reproducibility?

Key findings

  • Prompt design (PD) is the most prevalent paradigm, appearing in 78 studies.
  • A total of 114 studies were reviewed from 2022–2024.
  • In 12 papers, PD, PL, and PT terms were used interchangeably.
  • ChatGPT is the most commonly used LLM, with seven papers using it for processing sensitive clinical data.
  • Chain-of-Thought is the most common prompt engineering technique.
  • 64% of PD studies lack non-prompt-related baselines.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.