[Paper Review] Evaluating Verifiability in Generative Search Engines
The paper audits four commercial generative search engines for verifiability, finding high fluency but low citation recall (51.5%) and precision (74.5%), with accuracy issues impacting trust.
Generative search engines directly generate responses to user queries, along with in-line citations. A prerequisite trait of a trustworthy generative search engine is verifiability, i.e., systems should cite comprehensively (high citation recall; all statements are fully supported by citations) and accurately (high citation precision; every cite supports its associated statement). We conduct human evaluation to audit four popular generative search engines -- Bing Chat, NeevaAI, perplexity.ai, and YouChat -- across a diverse set of queries from a variety of sources (e.g., historical Google user queries, dynamically-collected open-ended questions on Reddit, etc.). We find that responses from existing generative search engines are fluent and appear informative, but frequently contain unsupported statements and inaccurate citations: on average, a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence. We believe that these results are concerningly low for systems that may serve as a primary tool for information-seeking users, especially given their facade of trustworthiness. We hope that our results further motivate the development of trustworthy generative search engines and help researchers and users better understand the shortcomings of existing commercial systems.
Motivation & Objective
- Define citation recall and citation precision as evaluation metrics for verifiability in generative search engines.
- Conduct a large-scale human evaluation of four commercial engines across diverse query distributions.
- Analyze how fluency, perceived utility, and verifiability interact in practice.
- Provide open annotations to support further research on trustworthy generative search engines.
Proposed method
- Define verification metrics: citation recall, citation precision, and citation F1.
- Segment each response into statements and associated citations to measure support.
- Use AIS (attributed to identified sources) judgments to assess whether statements are fully supported by citations.
- Evaluate fluency and perceived utility via annotator judgments on a 5-point Likert scale.
- Assess 12 query distributions with 1450 total queries across four engines.
- Release annotation data to facilitate reproducibility.
Experimental results
Research questions
- RQ1What are the levels of citation recall and citation precision across popular generative search engines?
- RQ2How do fluency and perceived utility relate to verifiability metrics in practice?
- RQ3Do systems exhibit trade-offs between recall and precision, and how does this affect user perception?
- RQ4Is higher citation precision associated with higher similarity to cited sources, and how does that relate to perceived utility?
Key findings
- Across engines, only 51.5% of generated sentences are fully supported by citations (recall).
- Only 74.5% of citations fully support their associated statements (precision).
- Perceive utility is inversely correlated with citation precision (r = -0.96).
- Perplexity.ai achieves the highest average citation recall (68.7), while Bing Chat achieves the highest average precision (89.5).
- Bing Chat often copies text from sources, yielding high precision but lower perceived utility due to irrelevance.
- YouChat shows low citation precision but higher perceived utility, illustrating the trade-off between faithfulness and usefulness.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.