Skip to main content
QUICK REVIEW

[论文解读] Exploring Qualitative Research Using LLMs

Muneera Bano, Didar Zowghi|arXiv (Cornell University)|Jun 23, 2023
Computational and Text Analysis Methods被引用 11
一句话总结

本论文比较人类与大语言模型在 Alexa 应用评审上的分类与推理,发现部分对齐以及人机协同的潜力。

ABSTRACT

The advent of AI driven large language models (LLMs) have stirred discussions about their role in qualitative research. Some view these as tools to enrich human understanding, while others perceive them as threats to the core values of the discipline. This study aimed to compare and contrast the comprehension capabilities of humans and LLMs. We conducted an experiment with small sample of Alexa app reviews, initially classified by a human analyst. LLMs were then asked to classify these reviews and provide the reasoning behind each classification. We compared the results with human classification and reasoning. The research indicated a significant alignment between human and ChatGPT 3.5 classifications in one third of cases, and a slightly lower alignment with GPT4 in over a quarter of cases. The two AI models showed a higher alignment, observed in more than half of the instances. However, a consensus across all three methods was seen only in about one fifth of the classifications. In the comparison of human and LLMs reasoning, it appears that human analysts lean heavily on their individual experiences. As expected, LLMs, on the other hand, base their reasoning on the specific word choices found in app reviews and the functional components of the app itself. Our results highlight the potential for effective human LLM collaboration, suggesting a synergistic rather than competitive relationship. Researchers must continuously evaluate LLMs role in their work, thereby fostering a future where AI and humans jointly enrich qualitative research.

研究动机与目标

  • 激发对 AI 驱动的 LLM 在定性研究中作用的理解。
  • 评估与人类分析师相比,LLMs 在定性数据分类上的表现有多好。
  • 调查人类与 LLM 在定性分类中的推理过程。
  • 探讨在定性研究中人与 LLM 之间有效协作的潜力。

提出的方法

  • 在一小部分由人类分析师最初分类的 Alexa 应用评审样本上进行实验。
  • 要求 LLM 对评审进行分类并提供每次分类背后的推理。
  • 将 LLM 的分类与推理与人类分类和人类推理进行比较。
  • 衡量人类、ChatGPT 3.5 与 GPT-4 分类之间的对齐程度。
  • 分析人类与 LLM 之间推理风格的差异。

实验结果

研究问题

  • RQ1LLM 的分类与人类在 Alexa 应用评审上的分类有多一致?
  • RQ2在对定性数据进行分类时,人类与 LLM 的推理过程有何比较?
  • RQ3在人类、ChatGPT 3.5 和 GPT-4 的分类之间的一致性水平是多少?
  • RQ4人类和 LLM 是否展示出互补的优势,暗示协作潜力?

主要发现

  • 大约三分之一的分类在人类与 ChatGPT 3.5 之间对齐。
  • 人类与 GPT-4 之间的对齐稍低(超过四分之一)。
  • 这两种 AI 模型彼此之间的对齐更高(在超过一半的实例中)。
  • 三者之间达成一致大约出现在五分之一的分类中。
  • 人类倾向于依赖个人经验,而 LLM 的推理基于单词选择和应用组件。
  • 结果表明在定性研究中存在人机协同的潜力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。