Skip to main content
QUICK REVIEW

[논문 리뷰] Evaluating Verifiability in Generative Search Engines

Nelson F. Liu, Tianyi Zhang|arXiv (Cornell University)|2023. 04. 19.
Misinformation and Its Impacts인용 수 14
한 줄 요약

본 논문은 네 가지 상용 생성형 검색 엔진의 검증 가능성을 평가하고, 높은 유창성을 발견했으나 인용 재현률(51.5%)과 정밀도(74.5%)가 낮아 신뢰도에 영향을 주는 정확도 문제를 확인했다.

ABSTRACT

Generative search engines directly generate responses to user queries, along with in-line citations. A prerequisite trait of a trustworthy generative search engine is verifiability, i.e., systems should cite comprehensively (high citation recall; all statements are fully supported by citations) and accurately (high citation precision; every cite supports its associated statement). We conduct human evaluation to audit four popular generative search engines -- Bing Chat, NeevaAI, perplexity.ai, and YouChat -- across a diverse set of queries from a variety of sources (e.g., historical Google user queries, dynamically-collected open-ended questions on Reddit, etc.). We find that responses from existing generative search engines are fluent and appear informative, but frequently contain unsupported statements and inaccurate citations: on average, a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence. We believe that these results are concerningly low for systems that may serve as a primary tool for information-seeking users, especially given their facade of trustworthiness. We hope that our results further motivate the development of trustworthy generative search engines and help researchers and users better understand the shortcomings of existing commercial systems.

연구 동기 및 목표

  • 생성형 검색 엔진에서 검증 가능성을 평가하기 위한 지표로 인용 재현률(citation recall)과 인용 정밀도(citation precision)를 정의한다.
  • 다양한 쿼리 분포에 걸쳐 네 가지 상용 엔진에 대한 대규모 인간 평가를 수행한다.
  • 유창성, 인지된 유용성, 및 검증 가능성이 실제로 어떻게 상호 작용하는지 분석한다.
  • 신뢰할 수 있는 생성형 검색 엔진에 대한 향후 연구를 지원하기 위해 공개 주석을 제공한다.

제안 방법

  • 검증 지표를 정의한다: 인용 재현률(citation recall), 인용 정밀도(citation precision), 및 인용 F1.
  • 각 응답을 진술과 관련 인용으로 분할하여 지지 여부를 측정한다.
  • 식별된 출처에 귀속된 AIS 판단을 사용하여 진술이 인용으로 완전히 뒷받침되는지 평가한다.
  • 5점 리커트 척도에서 주석가의 판단을 통해 유창성과 인지된 유용성을 평가한다.
  • 네 가지 엔진에 걸친 총 1450개의 쿼리로 12개 쿼리 분포를 평가한다.
  • 재현성을 촉진하기 위해 주석 데이터를 공개한다.

실험 결과

연구 질문

  • RQ1인기 있는 생성형 검색 엔진 전반에서 인용 재현률과 인용 정밀도의 수준은 어떠한가?
  • RQ2유창성과 인지된 유용성은 실제로 검증 가능성 지표와 어떤 관계가 있는가?
  • RQ3시스템 간 재현률과 정밀도 간의 트레이드오프가 나타나며, 이것이 사용자의 인식에 어떻게 영향을 미치는가?
  • RQ4높은 인용 정밀도가 인용된 소스와의 더 높은 유사성과 관련이 있으며, 이것이 인지된 유용성과 어떻게 연관되는가?

주요 결과

  • 엔진 전반에서 생성된 문장의 51.5%만이 인용으로 완전히 뒷받침된다(재현률).
  • 인용의 74.5%만이 관련 진술을 완전히 뒷받침한다(정밀도).
  • 인지된 유용성은 인용 정밀도와 부의 상관관계를 보이며(r = -0.96).
  • Perplexity.ai가 평균 인용 재현률(68.7)이 가장 높고, Bing Chat이 평균 정밀도(89.5)가 가장 높다.
  • Bing Chat은 종종 소스의 텍스트를 복사해 높은 정밀도를 보이지만 무관성으로 인해 인지된 유용성은 낮아진다.
  • YouChat은 인용 정밀도가 낮지만 인지된 유용성은 더 높아 신뢰도와 유용성 간의 트레이드오프를 보여준다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.