Skip to main content
QUICK REVIEW

[论文解读] Comparing Traditional and LLM-based Search for Consumer Choice: A Randomized Experiment

Sofia Eleni Spatharioti, David Rothschild|arXiv (Cornell University)|Jul 7, 2023
Consumer Market Behavior and Pricing被引用 25
一句话总结

本研究通过随机化实验比较传统搜索与基于LLM的搜索工具,发现使用LLMs可更快完成任务且满意度更高,但指出若不通过高亮等缓解措施容易过度依赖错误信息的风险。

ABSTRACT

Recent advances in the development of large language models are rapidly changing how online applications function. LLM-based search tools, for instance, offer a natural language interface that can accommodate complex queries and provide detailed, direct responses. At the same time, there have been concerns about the veracity of the information provided by LLM-based tools due to potential mistakes or fabrications that can arise in algorithmically generated text. In a set of online experiments we investigate how LLM-based search changes people's behavior relative to traditional search, and what can be done to mitigate overreliance on LLM-based output. Participants in our experiments were asked to solve a series of decision tasks that involved researching and comparing different products, and were randomly assigned to do so with either an LLM-based search tool or a traditional search engine. In our first experiment, we find that participants using the LLM-based tool were able to complete their tasks more quickly, using fewer but more complex queries than those who used traditional search. Moreover, these participants reported a more satisfying experience with the LLM-based search tool. When the information presented by the LLM was reliable, participants using the tool made decisions with a comparable level of accuracy to those using traditional search, however we observed overreliance on incorrect information when the LLM erred. Our second experiment further investigated this issue by randomly assigning some users to see a simple color-coded highlighting scheme to alert them to potentially incorrect or misleading information in the LLM responses. Overall we find that this confidence-based highlighting substantially increases the rate at which users spot incorrect information, improving the accuracy of their overall decisions while leaving most other measures unaffected.

研究动机与目标

  • 评估基于LLM的搜索在消费者决策任务中与传统搜索相比对用户行为的影响。
  • 评估在基于LLM的搜索与传统搜索下的任务时长与查询特征。
  • 在LLM信息可靠时与包含错误时,评估决策准确性。
  • 测试一种简单的错误警示/高亮技术以缓解对LLM输出的过度依赖。

提出的方法

  • 进行线上随机化实验,让参与者解决产品对比任务。
  • 将参与者分配到基于LLM的搜索或传统搜索条件。
  • 测量任务完成时间、查询复杂度与用户满意度。
  • 根据LLM信息的可靠性评估决策准确性。
  • 引入颜色编码的高亮以标记潜在不正确的信息并评估其效果。

实验结果

研究问题

  • RQ1基于LLM的搜索是否比传统搜索能降低任务完成时间?
  • RQ2在信息可靠时,用户对LLM输出的依赖是否比对传统搜索更高或更低?
  • RQ3警示/高亮机制是否通过减少对LLM的过度依赖来提高准确性?
  • RQ4两种搜索模式下用户满意度和查询特征有何差异?

主要发现

  • 基于LLM的搜索使任务完成速度比传统搜索更快。
  • 使用LLMs的参与者使用的查询更少但更复杂。
  • 当LLM信息可靠时,决策准确性可与传统搜索相媲美。
  • 当LLM出错时,LLM用户对不正确信息的过度依赖更明显。
  • 颜色编码的高亮以标记潜在不正确信息可以提高错误检测并改善决策准确性。
  • 高亮方法对其他指标的影响有限。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。