Skip to main content
QUICK REVIEW

[论文解读] Large Language Models for Large-Scale, Rigorous Qualitative Analysis in Applied Health Services Research

Sasha Ronaghi, Emma‐Louise Aveling|arXiv (Cornell University)|Jan 20, 2026
Health Policy Implementation Science被引用 0
一句话总结

论文提出一个面向任务的人工–LLM框架,将大型语言模型应用于多地点定性健康服务研究,在两个任务中在保持分析严谨性的同时实现效率提升:定性综合用于反馈报告与演绎编码用于改进干预。

ABSTRACT

Large language models (LLMs) show promise for improving the efficiency of qualitative analysis in large, multi-site health-services research. Yet methodological guidance for LLM integration into qualitative analysis and evidence of their impact on real-world research methods and outcomes remain limited. We developed a model- and task-agnostic framework for designing human-LLM qualitative analysis methods to support diverse analytic aims. Within a multi-site study of diabetes care at Federally Qualified Health Centers (FQHCs), we leveraged the framework to implement human-LLM methods for (1) qualitative synthesis of researcher-generated summaries to produce comparative feedback reports and (2) deductive coding of 167 interview transcripts to refine a practice-transformation intervention. LLM assistance enabled timely feedback to practitioners and the incorporation of large-scale qualitative data to inform theory and practice changes. This work demonstrates how LLMs can be integrated into applied health-services research to enhance efficiency while preserving rigor, offering guidance for continued innovation with LLMs in qualitative research.

研究动机与目标

  • 为应用于健康服务研究的定性分析方法开发一个面向任务与模型无关的人工–LLM框架。
  • 通过覆盖12个联邦认证的卫生中心(FQHCs)的多地点糖尿病护理研究来演示框架。
  • 使用LLMs来(a)生成比较性站点级反馈报告,以及(b)执行演绎编码以完善一个实践转型干预。
  • 评估LLM辅助分析在现实研究中对效率、严谨性和解释控制的影响。

提出的方法

  • 定义任务:在小样本数据上明确目标、产出与研究人员参与度。
  • 设计人–LLM方法:分解任务、明确目的、并为每个部分测试不同的人/ AI 配置。
  • 在小规模数据上评估方法:比较有无LLMs的输出,使用特定任务的严谨性标准(基础证据 grounding、理论-数据整合、相关性)。
  • 将方法应用并在更大数据集上评估完整任务,以评估效率和对研究目标的影响。
  • 任务1:定性综合,跨22个护理领域生成比较性站点级摘要,使用少量示例提示与领域定义将站点数据组织成主题,并通过LLM协助实现横向综合。
  • 任务2:演绎性定性编码以改进糖尿病护理干预,使用检索增强生成(RAG),结合基于嵌入的检索与结构化子问题方法,生成带引语的编码输出。
Figure 1: The framework we developed and applied for developing task-specific human-LLM qualitative analysis methods.
Figure 1: The framework we developed and applied for developing task-specific human-LLM qualitative analysis methods.

实验结果

研究问题

  • RQ1如何设计面向任务的人–LLM方法,在大规模定性健康服务研究中在提高效率的同时维持严谨性?
  • RQ2LLMs在多大程度上能够组织数据并生成支持但不替代研究者解释性分析的摘要或编码输出?
  • RQ3在多地点糖尿病护理研究中,LLM辅助定性分析对时间效率和对干预改进的影响如何?

主要发现

  • LLMs能够对站点级摘要进行主题化组织,在小规模测试中将起草比较性反馈报告的时间缩短了30–55%。
  • LLM辅助输出在综合任务的质量方面能够与手动定性组织相匹配,但需要研究者进行解释以确保可操作性并符合领域定义。
  • LLMs可以通过检索增强生成在大篇幅转录文本上支持演绎编码,但若研究者不进行验证和情境化,输出可能缺乏深度或背景信息。
  • 人–LLM协作能够保持解释控制,研究者保留最终判断,确保输出仍然 anchored 于数据与分析目标。
  • 该框架实现了将167份转录文本高效纳入19个实践领域代码,用于 eight 个额外站点的实施改进。
  • 研究强调需要原始数据访问并需要慎重设计,以在LLM生成的发现中维持透明度、可信度、自省性和偏差评估。
Figure 2: Illustrative differences in cross-site synthesis output by human and LLM (independently) for telehealth and appointment management themes within the “Information and Communication Technology” domain
Figure 2: Illustrative differences in cross-site synthesis output by human and LLM (independently) for telehealth and appointment management themes within the “Information and Communication Technology” domain

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。