[Paper Review] Citation: A Key to Building Responsible and Accountable Large Language Models
The paper argues for embedding a citation mechanism in LLMs to improve transparency, accountability, and IP/ethical governance, addressing both non-parametric and parametric content. It discusses implementation strategies, pitfalls, and a research agenda.
Large Language Models (LLMs) bring transformative benefits alongside unique challenges, including intellectual property (IP) and ethical concerns. This position paper explores a novel angle to mitigate these risks, drawing parallels between LLMs and established web systems. We identify "citation" - the acknowledgement or reference to a source or evidence - as a crucial yet missing component in LLMs. Incorporating citation could enhance content transparency and verifiability, thereby confronting the IP and ethical issues in the deployment of LLMs. We further propose that a comprehensive citation mechanism for LLMs should account for both non-parametric and parametric content. Despite the complexity of implementing such a citation mechanism, along with the potential pitfalls, we advocate for its development. Building on this foundation, we outline several research problems in this area, aiming to guide future explorations towards building more responsible and accountable LLMs.
Motivation & Objective
- Motivate the need to manage IP and ethical concerns in LLMs by drawing parallels to the Web and search engines.
- Introduce the concept of citations as a missing but crucial component in LLMs to enhance transparency and accountability.
- Outline potential strategies for citing non-parametric content and for attributing parametric content.
- Identify major challenges and research questions to guide future development of responsible LLMs.
Proposed method
- Define when and how citations should be used in LLM outputs (pre-hoc and post-hoc for non-parametric content).
- Propose a hybrid system combining LLMs with information retrieval to supply non-parametric citations.
- Discuss the idea of source identifiers or tokens to trace parametric content back to sources.
- Survey potential pitfalls and barriers, including hallucination, citation bias, and legal implications.
Experimental results
Research questions
- RQ1When should LLMs provide citations for generated information?
- RQ2How can LLMs cite non-parametric content via retrieval-augmented methods?
- RQ3How can LLMs attribute parametric content to training data sources?
- RQ4What are the major pitfalls in citation mechanisms for LLMs and how can they be mitigated?
- RQ5What research problems must be solved to enable reliable, up-to-date, and unbiased citations in LLMs?
Key findings
- Citations could address IP and ethical concerns by signaling sources and enabling verification.
- Non-parametric content can be cited via pre-hoc retrieval or post-hoc insertion, potentially combined for robustness.
- Parametric content poses attribution challenges due to high-dimensional training representations; source identifiers are a potential solution.
- Citations introduce risks such as over-citation, inaccurate or outdated sources, misinformation propagation, bias, and reduced creativity.
- A broad research agenda is outlined to tackle timing, accuracy, reliability, bias, and legal issues in citation-enabled LLMs.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.