[Paper Review] Comparing Traditional and LLM-based Search for Consumer Choice: A Randomized Experiment
The study compares traditional search with an LLM-based search tool through randomized experiments, finding faster task completion and higher satisfaction with LLMs, but noting risks of overreliance on incorrect information unless mitigated by highlighting.
Recent advances in the development of large language models are rapidly changing how online applications function. LLM-based search tools, for instance, offer a natural language interface that can accommodate complex queries and provide detailed, direct responses. At the same time, there have been concerns about the veracity of the information provided by LLM-based tools due to potential mistakes or fabrications that can arise in algorithmically generated text. In a set of online experiments we investigate how LLM-based search changes people's behavior relative to traditional search, and what can be done to mitigate overreliance on LLM-based output. Participants in our experiments were asked to solve a series of decision tasks that involved researching and comparing different products, and were randomly assigned to do so with either an LLM-based search tool or a traditional search engine. In our first experiment, we find that participants using the LLM-based tool were able to complete their tasks more quickly, using fewer but more complex queries than those who used traditional search. Moreover, these participants reported a more satisfying experience with the LLM-based search tool. When the information presented by the LLM was reliable, participants using the tool made decisions with a comparable level of accuracy to those using traditional search, however we observed overreliance on incorrect information when the LLM erred. Our second experiment further investigated this issue by randomly assigning some users to see a simple color-coded highlighting scheme to alert them to potentially incorrect or misleading information in the LLM responses. Overall we find that this confidence-based highlighting substantially increases the rate at which users spot incorrect information, improving the accuracy of their overall decisions while leaving most other measures unaffected.
Motivation & Objective
- Assess how LLM-based search alters user behavior compared with traditional search in consumer decision tasks.
- Evaluate task duration and query characteristics under LLM-based versus traditional search.
- Examine decision accuracy when LLM responses are reliable versus when they contain errors.
- Test a simple error-alerting technique to mitigate overreliance on LLM outputs.
Proposed method
- Conduct online randomized experiments with participants solving product comparison tasks.
- Assign participants to LLM-based search or traditional search conditions.
- Measure task completion time, query complexity, and user satisfaction.
- Assess decision accuracy relative to information reliability of the LLM.
- Introduce color-coded highlighting to flag potentially incorrect information and evaluate its effect.
Experimental results
Research questions
- RQ1Does LLM-based search reduce task completion time compared to traditional search?
- RQ2Do users rely on LLM outputs more or less than on traditional search when information is reliable?
- RQ3Can an alerting/highlighting mechanism improve accuracy by reducing overreliance on LLMs?
- RQ4How do user satisfaction and query characteristics differ between search modalities?
Key findings
- LLM-based search enables faster task completion than traditional search.
- Participants using LLMs use fewer but more complex queries.
- When LLM information is reliable, decision accuracy is comparable to traditional search.
- LLM users show overreliance on incorrect information when the LLM errs.
- Color-coded highlighting to flag potentially incorrect information increases error detection and improves decision accuracy.
- The highlighting method has limited impact on other measures.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.