[Paper Review] Collecting Qualitative Data at Scale with Large Language Models: A Case Study
This study evaluates large language model (LLM)-augmented chatbots for scalable qualitative data collection, demonstrating that LLM-powered dynamic probing and conversation summarization significantly improve response richness and user satisfaction over rule-based chatbots, while users still prefer human interviewers despite improved AI interactions.
Chatbots have shown promise as tools to scale qualitative data collection. Recent advances in Large Language Models (LLMs) could accelerate this process by allowing researchers to easily deploy sophisticated interviewing chatbots. We test this assumption by conducting a large-scale user study (n=399) evaluating 3 different chatbots, two of which are LLM-based and a baseline which employs hard-coded questions. We evaluate the results with respect to participant engagement and experience, established metrics of chatbot quality grounded in theories of effective communication, and a novel scale evaluating "richness" or the extent to which responses capture the complexity and specificity of the social context under study. We find that, while the chatbots were able to elicit high-quality responses based on established evaluation metrics, the responses rarely capture participants' specific motives or personalized examples, and thus perform poorly with respect to richness. We further find low inter-rater reliability between LLMs and humans in the assessment of both quality and richness metrics. Our study offers a cautionary tale for scaling and evaluating qualitative research with LLMs.
Motivation & Objective
- To address the scalability-qualitativeness trade-off in HCI research by leveraging LLMs to enhance chatbot-based data collection.
- To narrow the 'gulf of expectations' between users' perceptions of conversational AI and actual system capabilities.
- To evaluate whether LLM-augmented chatbots outperform rule-based chatbots in user engagement, response quality, and user experience.
- To explore the feasibility of using LLM-generated conversation summaries for scalable member checking in qualitative research.
- To provide an open-source framework for researchers to adopt and extend LLM-augmented chatbot designs.
Proposed method
- Developed an LLM-augmented chatbot capable of dynamically generating follow-up questions based on user input.
- Implemented a real-time conversation summarization module using LLMs to generate and present summaries for user confirmation.
- Designed a user study with 399 participants randomly assigned to interact with a rule-based chatbot or one of two LLM-augmented variants.
- Used both quantitative metrics (e.g., response length, engagement time) and qualitative feedback to assess user experience and response quality.
- Applied prompt engineering to enable dynamic probing and summary generation, minimizing reliance on rigid rule-based logic.
- Provided an open-source codebase and prompts to support replication and extension by other HCI researchers.

Experimental results
Research questions
- RQ1Does the integration of LLMs into chatbot-based data collection improve user engagement compared to rule-based systems?
- RQ2To what extent do LLM-augmented chatbots enhance the richness and depth of qualitative responses?
- RQ3How do users' expectations of conversational AI compare to the actual performance of LLM-powered chatbots?
- RQ4Can LLM-generated conversation summaries serve as an effective method for scalable member checking in qualitative research?
- RQ5How do users perceive LLM-augmented chatbots in comparison to traditional web surveys or human interviewers?
Key findings
- The LLM-augmented chatbot with dynamic probing significantly improved both quantitative and qualitative measures of user experience, indicating a narrowing of the expectations gap.
- Nearly 95% of participants agreed with the LLM-generated conversation summaries, demonstrating strong potential for scalable member checking.
- While engagement levels were similar across conditions, response richness was significantly higher in the LLM-augmented conditions.
- Users preferred the LLM-augmented chatbot over traditional web surveys, though they still expressed a preference for human interviewers.
- Participants reported expectations of personalization and the ability to ask clarifying questions, suggesting these features could further improve perceived quality.
- The study confirms that LLM-augmented chatbots can collect high-quality qualitative data at scale, offering a viable alternative to in-depth interviews in time-sensitive research contexts.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.