[Paper Review] From Prompt Engineering to Prompt Science With Human in the Loop
The paper proposes a four-phase, human-in-the-loop methodology inspired by qualitative coding to turn ad-hoc prompt engineering into verifiable, replicable prompt science for LLM-assisted research.
As LLMs make their way into many aspects of our lives, one place that warrants increased scrutiny with LLM usage is scientific research. Using LLMs for generating or analyzing data for research purposes is gaining popularity. But when such application is marred with ad-hoc decisions and engineering solutions, we need to be concerned about how it may affect that research, its findings, or any future works based on that research. We need a more scientific approach to using LLMs in our research. While there are several active efforts to support more systematic construction of prompts, they are often focused more on achieving desirable outcomes rather than producing replicable and generalizable knowledge with sufficient transparency, objectivity, or rigor. This article presents a new methodology inspired by codebook construction through qualitative methods to address that. Using humans in the loop and a multi-phase verification processes, this methodology lays a foundation for more systematic, objective, and trustworthy way of applying LLMs for analyzing data. Specifically, we show how a set of researchers can work through a rigorous process of labeling, deliberating, and documenting to remove subjectivity and bring transparency and replicability to prompt generation process. A set of experiments are presented to show how this methodology can be put in practice.
Motivation & Objective
- Motivate the need for scientific rigor when using LLMs in research and identify risks of ad-hoc prompt engineering.
- Introduce a systematic, transparent process to develop prompts and assess LLM outputs.
- Adapt qualitative coding with multiple assessors to create a replicable prompt construction codebook.
- Provide a multi-phase pipeline that ensures reliability, generalizability, and verifiability of prompts and responses.
Proposed method
- Adopt a codebook-building approach from qualitative coding to construct prompts.
- Implement a four-phase pipeline (set up, criteria establishment with ICR, iterative prompt development, validation) with human-in-the-loop assessment.
- Require at least two qualified researchers to participate and compute inter-coder reliability (ICR) such as Cohen’s kappa or Krippendorff’s alpha.
- Iteratively revise the codebook (criteria) and the prompt based on assessor disagreements to improve agreement and generalizability.
- Optionally validate the entire pipeline on a test data subset and compute ICR for the final assessment.
Experimental results
Research questions
- RQ1How can prompt generation for LLMs be made verifiably reliable and replicable across datasets, models, and times?
- RQ2What role do human assessors and codebook-like criteria play in achieving objective, transparent prompt generation?
- RQ3Can a multi-phase, qualitative-coding-inspired process reduce subjectivity and bias in LLM-driven data labeling or analysis?
- RQ4What are the costs and benefits of implementing prompt science versus traditional prompt engineering?
Key findings
- A multi-phase prompt construction process with human-in-the-loop yields more transparent, verifiable, and replicable prompts.
- Involvement of multiple researchers and formal ICR measures reduces individual biases and improves consistency in assessments.
- Documenting deliberations and decisions enhances openness and reproducibility for future researchers.
- Compared to ad-hoc prompt engineering, the proposed approach increases quality and understanding, albeit with higher costs.
- Optional validation phase can further ensure pipeline reliability across data samples.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.