Skip to main content
QUICK REVIEW

[Paper Review] Large Language Models for Large-Scale, Rigorous Qualitative Analysis in Applied Health Services Research

Sasha Ronaghi, Emma‐Louise Aveling|arXiv (Cornell University)|Jan 20, 2026
Health Policy Implementation Science0 citations
TL;DR

The paper presents a task-specific human–LLM framework to integrate large language models into multi-site qualitative health-services research, demonstrating efficiency gains while preserving analytic rigor in two tasks: qualitative synthesis for feedback reports and deductive coding to refine an intervention.

ABSTRACT

Large language models (LLMs) show promise for improving the efficiency of qualitative analysis in large, multi-site health-services research. Yet methodological guidance for LLM integration into qualitative analysis and evidence of their impact on real-world research methods and outcomes remain limited. We developed a model- and task-agnostic framework for designing human-LLM qualitative analysis methods to support diverse analytic aims. Within a multi-site study of diabetes care at Federally Qualified Health Centers (FQHCs), we leveraged the framework to implement human-LLM methods for (1) qualitative synthesis of researcher-generated summaries to produce comparative feedback reports and (2) deductive coding of 167 interview transcripts to refine a practice-transformation intervention. LLM assistance enabled timely feedback to practitioners and the incorporation of large-scale qualitative data to inform theory and practice changes. This work demonstrates how LLMs can be integrated into applied health-services research to enhance efficiency while preserving rigor, offering guidance for continued innovation with LLMs in qualitative research.

Motivation & Objective

  • Develop a task- and model-agnostic framework for human–LLM qualitative analysis methods in applied health services research.
  • Demonstrate framework via a multi-site diabetes care study across 12 Federally Qualified Health Centers (FQHCs).
  • Use LLMs to (a) generate comparative site-level feedback reports and (b) perform deductive coding to refine a practice-transformation intervention.
  • Assess how LLM-assisted analysis affects efficiency, rigor, and interpretive control in real-world research.

Proposed method

  • Define Task: clarify goals, outputs, and required researcher involvement on small data samples.
  • Design human–LLM method: decompose tasks, specify purposes, and test different human/AI configurations for each part.
  • Evaluate method on small-scale data: compare outputs with and without LLMs using task-specific rigor criteria (grounding, theory-data integration, relevance).
  • Apply and evaluate method to full tasks on the larger dataset to assess efficiency and impact on research goals.
  • Task 1: qualitative synthesis to produce comparative site-level summaries across 22 care domains, using few-shot prompts and domain definitions to organize site data into themes with LLM assistance for cross-site syntheses.
  • Task 2: deductive qualitative coding to refine a diabetes care intervention, using retrieval-augmented generation (RAG) with embedding-based retrieval and a structured sub-question approach to generate coded outputs with quotes.
Figure 1: The framework we developed and applied for developing task-specific human-LLM qualitative analysis methods.
Figure 1: The framework we developed and applied for developing task-specific human-LLM qualitative analysis methods.

Experimental results

Research questions

  • RQ1How can task-specific human–LLM methods be designed to preserve rigor while improving efficiency in large-scale qualitative health-services research?
  • RQ2To what extent can LLMs organize data and generate summaries or coding outputs that support but do not replace researchers' interpretive analysis?
  • RQ3What are the practical impacts of LLM-assisted qualitative analysis on time efficiency and on informing intervention refinement in a multi-site diabetes care study?

Key findings

  • LLMs can organize site-level summaries thematically, reducing time to draft comparative feedback reports by up to 30–55% in small-scale tests.
  • LLM-assisted outputs can match manual qualitative organization in quality for synthesis tasks but require researcher interpretation to ensure actionability and alignment with domain definitions.
  • LLMs can support deductive coding across large transcripts via retrieval-augmented generation, but outputs may lack depth or context unless researchers validate and contextualize them.
  • Human–LLM collaboration preserves interpretive control, with researchers retaining final judgments, ensuring outputs remain anchored in data and analytic goals.
  • The framework enables efficient incorporation of 167 transcripts into 19 practice-area codes, informing intervention refinement for implementation in eight additional sites.
  • The study highlights the need for raw data access and careful design to maintain transparency, credibility, reflexivity, and bias assessment in LLM-generated findings.
Figure 2: Illustrative differences in cross-site synthesis output by human and LLM (independently) for telehealth and appointment management themes within the “Information and Communication Technology” domain
Figure 2: Illustrative differences in cross-site synthesis output by human and LLM (independently) for telehealth and appointment management themes within the “Information and Communication Technology” domain

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.