Skip to main content
QUICK REVIEW

[Paper Review] Exploring Qualitative Research Using LLMs

Muneera Bano, Didar Zowghi|arXiv (Cornell University)|Jun 23, 2023
Computational and Text Analysis Methods11 citations
TL;DR

The paper compares human and LLM classifications and reasoning on Alexa app reviews, finding partial alignment and potential for synergistic human-LLM collaboration.

ABSTRACT

The advent of AI driven large language models (LLMs) have stirred discussions about their role in qualitative research. Some view these as tools to enrich human understanding, while others perceive them as threats to the core values of the discipline. This study aimed to compare and contrast the comprehension capabilities of humans and LLMs. We conducted an experiment with small sample of Alexa app reviews, initially classified by a human analyst. LLMs were then asked to classify these reviews and provide the reasoning behind each classification. We compared the results with human classification and reasoning. The research indicated a significant alignment between human and ChatGPT 3.5 classifications in one third of cases, and a slightly lower alignment with GPT4 in over a quarter of cases. The two AI models showed a higher alignment, observed in more than half of the instances. However, a consensus across all three methods was seen only in about one fifth of the classifications. In the comparison of human and LLMs reasoning, it appears that human analysts lean heavily on their individual experiences. As expected, LLMs, on the other hand, base their reasoning on the specific word choices found in app reviews and the functional components of the app itself. Our results highlight the potential for effective human LLM collaboration, suggesting a synergistic rather than competitive relationship. Researchers must continuously evaluate LLMs role in their work, thereby fostering a future where AI and humans jointly enrich qualitative research.

Motivation & Objective

  • Motivate understanding of the role of AI-driven LLMs in qualitative research.
  • Assess how well LLMs can classify qualitative data compared with human analysts.
  • Investigate the reasoning processes of humans and LLMs in qualitative classification.
  • Explore the potential for effective collaboration between humans and LLMs in qualitative studies.

Proposed method

  • Conducted an experiment on a small sample of Alexa app reviews initially classified by a human analyst.
  • Asked LLMs to classify the reviews and provide the reasoning behind each classification.
  • Compared LLM classifications and reasoning with human classifications and human reasoning.
  • Measured alignment across human, ChatGPT 3.5, and GPT-4 classifications.
  • Analyzed differences in reasoning styles between humans and LLMs.

Experimental results

Research questions

  • RQ1How aligned are LLM classifications with human classifications on Alexa app reviews?
  • RQ2How do the reasoning processes of humans and LLMs compare when classifying qualitative data?
  • RQ3What is the level of agreement among human, ChatGPT 3.5, and GPT-4 classifications?
  • RQ4Do humans and LLMs exhibit complementary strengths suggesting collaborative potential?

Key findings

  • About one third of classifications were aligned between humans and ChatGPT 3.5.
  • Slightly lower alignment between humans and GPT-4 (over a quarter).
  • The two AI models showed higher alignment with each other (more than half of the instances).
  • Consensus across all three methods occurred in about one fifth of classifications.
  • Humans tend to rely on individual experience, while LLMs base reasoning on word choices and app components.
  • Results suggest a synergistic potential for human-LLM collaboration in qualitative research.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.