Skip to main content
QUICK REVIEW

[논문 리뷰] Comparing Traditional and LLM-based Search for Consumer Choice: A Randomized Experiment

Sofia Eleni Spatharioti, David Rothschild|arXiv (Cornell University)|2023. 07. 07.
Consumer Market Behavior and Pricing인용 수 25
한 줄 요약

연구는 전통 검색과 LLM 기반 검색 도구를 무작위 실험으로 비교하여 더 빠른 작업 완료와 LLM에 대한 더 높은 만족도를 발견했지만, 하이라이팅으로 보완하지 않으면 부정확한 정보에 과도하게 의존할 위험이 있음

ABSTRACT

Recent advances in the development of large language models are rapidly changing how online applications function. LLM-based search tools, for instance, offer a natural language interface that can accommodate complex queries and provide detailed, direct responses. At the same time, there have been concerns about the veracity of the information provided by LLM-based tools due to potential mistakes or fabrications that can arise in algorithmically generated text. In a set of online experiments we investigate how LLM-based search changes people's behavior relative to traditional search, and what can be done to mitigate overreliance on LLM-based output. Participants in our experiments were asked to solve a series of decision tasks that involved researching and comparing different products, and were randomly assigned to do so with either an LLM-based search tool or a traditional search engine. In our first experiment, we find that participants using the LLM-based tool were able to complete their tasks more quickly, using fewer but more complex queries than those who used traditional search. Moreover, these participants reported a more satisfying experience with the LLM-based search tool. When the information presented by the LLM was reliable, participants using the tool made decisions with a comparable level of accuracy to those using traditional search, however we observed overreliance on incorrect information when the LLM erred. Our second experiment further investigated this issue by randomly assigning some users to see a simple color-coded highlighting scheme to alert them to potentially incorrect or misleading information in the LLM responses. Overall we find that this confidence-based highlighting substantially increases the rate at which users spot incorrect information, improving the accuracy of their overall decisions while leaving most other measures unaffected.

연구 동기 및 목표

  • 소비자 의사결정 과제에서 LLM 기반 검색이 전통 검색과 비교하여 사용자 행동을 어떻게 바꾸는지 평가한다.
  • LLM 기반 검색과 전통 검색에서의 작업 지속 시간과 질의 특성을 평가한다.
  • LLM 응답이 신뢰할 수 있을 때와 오류가 있을 때의 의사결정 정확도를 검토한다.
  • LLM 출력에 대한 과도한 의존을 완화하기 위한 간단한 오류 경고 기법을 시험한다.

제안 방법

  • 참여자들이 제품 비교 과제를 해결하는 온라인 무작위 실험을 수행한다.
  • 참여자를 LLM 기반 검색 또는 전통 검색 조건에 배정한다.
  • 작업 완료 시간, 질의 복잡성, 그리고 사용자 만족도를 측정한다.
  • LLM의 정보 신뢰도에 비례한 의사결정 정확도를 평가한다.
  • 잠재적으로 잘못된 정보를 표시하기 위한 색상 코드 하이라이팅을 도입하고 그 효과를 평가한다.

실험 결과

연구 질문

  • RQ1LLM 기반 검색이 전통 검색에 비해 작업 완료 시간을 단축하는가?
  • RQ2정보가 신뢰할 수 있을 때 사용자들은 LLM 출력에 전통 검색보다 더 많이 의존하는가, 아니면 덜 의존하는가?
  • RQ3경고/하이라이트 메커니즘이 LLM에 대한 과도한 의존을 줄여 정확도를 향상시킬 수 있는가?
  • RQ4검색 방식 간에 사용자 만족도와 질의 특성이 어떻게 다른가?

주요 결과

  • LLM 기반 검색은 전통 검색보다 더 빠른 작업 완료를 가능하게 한다.
  • 참가자들이 LLM을 사용할 때 더 적고 더 복잡한 질의를 사용한다.
  • LLM 정보가 신뢰할 수 있을 때 의사결정 정확도는 전통 검색과 비슷하다.
  • LLM이 잘못될 때 LLM 사용자는 잘못된 정보에 과도하게 의존한다.
  • 잠재적으로 잘못된 정보를 표시하기 위한 색상 코드 하이라이팅은 오류 탐지를 증가시키고 의사결정 정확도를 향상시킨다.
  • 하이라이팅 방법은 다른 측정치에 제한된 영향을 미친다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.