[Paper Review] Large Language Models for Large-Scale, Rigorous Qualitative Analysis in Applied Health Services Research
The paper presents a task-specific human–LLM framework to integrate large language models into multi-site qualitative health-services research, demonstrating efficiency gains while preserving analytic rigor in two tasks: qualitative synthesis for feedback reports and deductive coding to refine an intervention.
Large language models (LLMs) show promise for improving the efficiency of qualitative analysis in large, multi-site health-services research. Yet methodological guidance for LLM integration into qualitative analysis and evidence of their impact on real-world research methods and outcomes remain limited. We developed a model- and task-agnostic framework for designing human-LLM qualitative analysis methods to support diverse analytic aims. Within a multi-site study of diabetes care at Federally Qualified Health Centers (FQHCs), we leveraged the framework to implement human-LLM methods for (1) qualitative synthesis of researcher-generated summaries to produce comparative feedback reports and (2) deductive coding of 167 interview transcripts to refine a practice-transformation intervention. LLM assistance enabled timely feedback to practitioners and the incorporation of large-scale qualitative data to inform theory and practice changes. This work demonstrates how LLMs can be integrated into applied health-services research to enhance efficiency while preserving rigor, offering guidance for continued innovation with LLMs in qualitative research.
Motivation & Objective
- Develop a task- and model-agnostic framework for human–LLM qualitative analysis methods in applied health services research.
- Demonstrate framework via a multi-site diabetes care study across 12 Federally Qualified Health Centers (FQHCs).
- Use LLMs to (a) generate comparative site-level feedback reports and (b) perform deductive coding to refine a practice-transformation intervention.
- Assess how LLM-assisted analysis affects efficiency, rigor, and interpretive control in real-world research.
Proposed method
- Define Task: clarify goals, outputs, and required researcher involvement on small data samples.
- Design human–LLM method: decompose tasks, specify purposes, and test different human/AI configurations for each part.
- Evaluate method on small-scale data: compare outputs with and without LLMs using task-specific rigor criteria (grounding, theory-data integration, relevance).
- Apply and evaluate method to full tasks on the larger dataset to assess efficiency and impact on research goals.
- Task 1: qualitative synthesis to produce comparative site-level summaries across 22 care domains, using few-shot prompts and domain definitions to organize site data into themes with LLM assistance for cross-site syntheses.
- Task 2: deductive qualitative coding to refine a diabetes care intervention, using retrieval-augmented generation (RAG) with embedding-based retrieval and a structured sub-question approach to generate coded outputs with quotes.

Experimental results
Research questions
- RQ1How can task-specific human–LLM methods be designed to preserve rigor while improving efficiency in large-scale qualitative health-services research?
- RQ2To what extent can LLMs organize data and generate summaries or coding outputs that support but do not replace researchers' interpretive analysis?
- RQ3What are the practical impacts of LLM-assisted qualitative analysis on time efficiency and on informing intervention refinement in a multi-site diabetes care study?
Key findings
- LLMs can organize site-level summaries thematically, reducing time to draft comparative feedback reports by up to 30–55% in small-scale tests.
- LLM-assisted outputs can match manual qualitative organization in quality for synthesis tasks but require researcher interpretation to ensure actionability and alignment with domain definitions.
- LLMs can support deductive coding across large transcripts via retrieval-augmented generation, but outputs may lack depth or context unless researchers validate and contextualize them.
- Human–LLM collaboration preserves interpretive control, with researchers retaining final judgments, ensuring outputs remain anchored in data and analytic goals.
- The framework enables efficient incorporation of 167 transcripts into 19 practice-area codes, informing intervention refinement for implementation in eight additional sites.
- The study highlights the need for raw data access and careful design to maintain transparency, credibility, reflexivity, and bias assessment in LLM-generated findings.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.