[Paper Review] CoQuest: Exploring Research Question Co-Creation with an LLM-based Agent
CoQuest is an LLM-based co-creation system that supports researchers in generating novel research questions through human-AI collaboration, using two interaction designs—breadth-first and depth-first RQ generation. Results from a 20-participant study show that while depth-first generation felt more creative during the task, breadth-first generation led to more creative and trustworthy outcomes overall, with AI processing delays enhancing user reflection and perceived control.
Developing novel research questions (RQs) often requires extensive literature reviews, especially in interdisciplinary fields. To support RQ development through human-AI co-creation, we leveraged Large Language Models (LLMs) to build an LLM-based agent system named CoQuest. We conducted an experiment with 20 HCI researchers to examine the impact of two interaction designs: breadth-first and depth-first RQ generation. The findings revealed that participants perceived the breadth-first approach as more creative and trustworthy upon task completion. Conversely, during the task, participants considered the depth-first generated RQs as more creative. Additionally, we discovered that AI processing delays allowed users to reflect on multiple RQs simultaneously, leading to a higher quantity of generated RQs and an enhanced sense of control. Our work makes both theoretical and practical contributions by proposing and evaluating a mental model for human-AI co-creation of RQs. We also address potential ethical issues, such as biases and over-reliance on AI, advocating for using the system to improve human research creativity rather than automating scientific inquiry.
Motivation & Objective
- To investigate how different AI interaction designs—breadth-first and depth-first—impact the co-creation of research questions in human-AI collaboration.
- To examine the role of AI processing delays in fostering user reflection, perceived control, and creativity during RQ generation.
- To evaluate user trust, perceived control, and cognitive biases in human-AI co-creation of research questions.
- To design and assess a system that integrates literature visualization, rationale explanation (AI Thoughts), and interactive feedback for improved co-creation experiences.
Proposed method
- CoQuest is a three-panel interface: RQ Flow Editor for RQ generation, Paper Graph Visualizer for literature exploration, and AI Thoughts panel for explaining AI’s reasoning.
- The system supports two interaction modes: breadth-first (generating multiple RQs early) and depth-first (focusing on one RQ at a time), with user feedback integrated iteratively.
- AI processing delays were intentionally introduced to allow users time to reflect and consider multiple RQs simultaneously.
- Users rated generated RQs on creativity and trust, and provided feedback to refine subsequent outputs.
- The system uses LLMs to generate RQs based on literature, with explanations (AI Thoughts) provided when users click on links between RQs.
- A within-subjects experimental design with 20 doctoral students compared the two interaction modes across multiple metrics, including perceived creativity, trust, and sense of control.

Experimental results
Research questions
- RQ1How does the breadth-first vs. depth-first interaction design influence perceived creativity and trust in AI-generated research questions?
- RQ2How do AI processing delays affect user reflection, perceived control, and the quality of co-created research questions?
- RQ3To what extent do cognitive biases, such as confirmation bias, affect user perception and evaluation of AI-generated RQs?
- RQ4How do explanations (AI Thoughts) and RQ rating features influence user engagement and reduce over-reliance on AI?
Key findings
- Participants rated the breadth-first approach as more creative and trustworthy upon task completion, despite initially finding the depth-first approach more creative during the task.
- AI processing delays enabled users to reflect on multiple RQs simultaneously, leading to a higher number of generated RQs and a stronger sense of perceived control.
- Users exhibited confirmation bias, with their sense of control decreasing when AI-generated RQs did not meet intrinsic expectations.
- The AI Thoughts feature was frequently used during wait times, enhancing user understanding and perceived control in the co-creation process.
- RQ rating and explanation features helped users actively evaluate AI outputs, reducing blind reliance and improving engagement.
- The system’s design effectively mitigated some risks of AI over-reliance, supporting active human thinking and better co-creation outcomes.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.