[论文解读] AI and the Problem of Knowledge Collapse
本文引入了“知识坍塌”这一概念——由于过度依赖大型语言模型(LLMs)而倾向于中心化、常见回应,导致社会对知识的理解趋于狭窄。通过模拟模型,研究发现,当AI生成内容的成本降低20%时,公众信念与真实情况的偏离程度会达到完全依赖全谱知识时的2.3倍,凸显了对创新和文化多样性的潜在威胁。
While artificial intelligence has the potential to process vast amounts of data, generate new insights, and unlock greater productivity, its widespread adoption may entail unforeseen consequences. We identify conditions under which AI, by reducing the cost of access to certain modes of knowledge, can paradoxically harm public understanding. While large language models are trained on vast amounts of diverse data, they naturally generate output towards the 'center' of the distribution. This is generally useful, but widespread reliance on recursive AI systems could lead to a process we define as "knowledge collapse", and argue this could harm innovation and the richness of human understanding and culture. However, unlike AI models that cannot choose what data they are trained on, humans may strategically seek out diverse forms of knowledge if they perceive them to be worthwhile. To investigate this, we provide a simple model in which a community of learners or innovators choose to use traditional methods or to rely on a discounted AI-assisted process and identify conditions under which knowledge collapse occurs. In our default model, a 20% discount on AI-generated content generates public beliefs 2.3 times further from the truth than when there is no discount. An empirical approach to measuring the distribution of LLM outputs is provided in theoretical terms and illustrated through a specific example comparing the diversity of outputs across different models and prompting styles. Finally, based on the results, we consider further research directions to counteract such outcomes.
研究动机与目标
- 探究广泛采用AI辅助知识获取是否可能因知识来源多样性减少而 paradoxically 降低公众理解水平。
- 建模个体或社群在AI成本折扣优势下,仍可能战略性地寻求多样化知识的条件。
- 实证测量不同LLM模型及提示策略下输出的知识多样性,尤其聚焦长尾知识。
- 识别因AI对信息的递归中介而对教育、创新和文化保存造成的系统性知识坍塌风险。
- 提出通过改进AI设计、提升透明度以及激励用户获取多样化知识来源的研究方向。
提出的方法
- 构建一个正向知识溢出模型,其中参与者可在使用带成本折扣的AI辅助内容与投资于全谱多样化知识来源之间进行选择。
- 采用模拟框架,建模社区中的信念形成过程,追踪AI成本折扣如何影响公众信念与真实知识分布之间的距离。
- 采用理论框架,通过不同提示下命名实体(如哲学家、思想流派)的频率分布来定义和测量LLM输出的多样性。
- 对比不同模型(GPT-3.5-turbo、Claude-3-sonnet、Gemini-Pro、Llama2-70b)在不同提示风格(简单查询 vs. 结构化请求多样化响应)下的LLM输出。
- 将频率分布截断为前600个最常见实体,以可视化并比较不同模型和提示下的响应集中程度。
- 利用来自2,693个已识别哲学实体的实证数据,评估提示风格如何影响长尾知识的呈现。
实验结果
研究问题
- RQ1在何种条件下,AI中介的知识获取会导致公众信念系统性地偏离真实知识分布?
- RQ2AI生成内容的成本折扣在多大程度上影响了个人与社群所获取知识的多样性与代表性?
- RQ3人类的战略性行为(如主动寻求非AI来源)在AI效率优势面前,能在多大程度上防止知识坍塌?
- RQ4不同提示策略在多大程度上影响LLM输出的多样性,特别是对代表性不足或长尾知识的呈现?
- RQ5知识坍塌对创新、文化保存以及多元世界观的公平获取具有何种影响?
主要发现
- AI生成内容成本降低20%时,公众信念与真实情况的距离相比无折扣情形扩大了2.3倍。
- LLMs系统性地偏好中心化、高频回应,导致输出高度集中于主导观点,忽视长尾知识。
- 明确要求多样化回应的提示策略(如列出特定地区的20位哲学家)可显著提升代表性不足实体的出现频率,降低少数名称的主导地位。
- 即使在GPT-3.5-turbo和Claude-3-sonnet等高质量模型中,输出分布默认仍高度倾斜,少数实体(如亚里士多德、康德)占据主导频率。
- 较少见实体(如约鲁巴哲学、豪萨哲学、阿维森纳)在各模型中的提及频率仍保持低位,表明其在AI输出中持续被边缘化。
- 当提示明确鼓励多样性(如提示v5)时,无单一实体占据主导地位,表明通过有意识的提示设计可显著提升输出多样性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。