Skip to main content
QUICK REVIEW

[论文解读] A Pilot Evaluation of ChatGPT and DALL-E 2 on Decision Making and Spatial Reasoning

Zhisheng Tang, Mayank Kejriwal|arXiv (Cornell University)|Feb 15, 2023
Explainable Artificial Intelligence (XAI)被引用 6
一句话总结

本试点研究使用非对抗性、中性提示,评估了ChatGPT和DALL-E 2在决策理性与空间推理方面的能力。DALL-E 2在每个空间推理提示下至少生成了一幅正确图像,尽管在理解物体指代方面表现良好,但常产生错误输出;ChatGPT在理性决策方面表现不一致,在许多情况下违反了冯·诺依曼-摩根斯特恩公理,尽管在复杂问题上推理正确。

ABSTRACT

We conduct a pilot study selectively evaluating the cognitive abilities (decision making and spatial reasoning) of two recently released generative transformer models, ChatGPT and DALL-E 2. Input prompts were constructed following neutral a priori guidelines, rather than adversarial intent. Post hoc qualitative analysis of the outputs shows that DALL-E 2 is able to generate at least one correct image for each spatial reasoning prompt, but most images generated are incorrect (even though the model seems to have a clear understanding of the objects mentioned in the prompt). Similarly, in evaluating ChatGPT on the rationality axioms developed under the classical Von Neumann-Morgenstern utility theorem, we find that, although it demonstrates some level of rational decision-making, many of its decisions violate at least one of the axioms even under reasonable constructions of preferences, bets, and decision-making prompts. ChatGPT's outputs on such problems generally tended to be unpredictable: even as it made irrational decisions (or employed an incorrect reasoning process) for some simpler decision-making problems, it was able to draw correct conclusions for more complex bet structures. We briefly comment on the nuances and challenges involved in scaling up such a 'cognitive' evaluation or conducting it with a closed set of answer keys ('ground truth'), given that these models are inherently generative and open-ended in responding to prompts.

研究动机与目标

  • 评估生成式AI模型(特别是ChatGPT和DALL-E 2)在决策与空间推理方面的认知能力。
  • 评估这些模型在冯·诺依曼-摩根斯特恩效用框架下是否遵循理性决策原则。
  • 检查在使用描述性提示进行空间推理任务时,生成输出的可靠性和一致性。
  • 探讨由于这些模型具有开放性、生成性特征,导致在扩展认知评估时面临的挑战。

提出的方法

  • 构建中性、非对抗性提示,以在无偏见或欺骗意图的情况下评估决策与空间推理能力。
  • 应用冯·诺依曼-摩根斯特恩理性公理(完备性、传递性、独立性与连续性)作为评估ChatGPT决策输出的基准。
  • 使用描述性、基于对象的提示,评估DALL-E 2生成空间构型准确视觉表征的能力。
  • 对模型输出进行事后定性分析,以评估其正确性、连贯性以及与预期推理或视觉结构的一致性。
  • 由于模型响应具有生成性和开放性特征,避免使用预定义的“真实答案”(“ground truth”)。
  • 在从简单到复杂的各种复杂度水平上,对两个模型进行了评估,涵盖决策与空间推理问题。

实验结果

研究问题

  • RQ1在可能存在空间关系误解的情况下,DALL-E 2在多大程度上能为空间推理提示生成准确的视觉表征?
  • RQ2ChatGPT在各种决策场景中,其对冯·诺依曼-摩根斯特恩效用理论理性公理的遵循程度有多一致?
  • RQ3为何ChatGPT的某些回答在复杂赌局中得出了正确结论,却在简单问题上表现出非理性推理?
  • RQ4由于这些模型具有开放性、非确定性响应模式,尝试扩展其认知评估时会面临哪些挑战?
  • RQ5缺乏固定‘真实答案’如何影响对生成式AI在认知任务上评估的可靠性和有效性?

主要发现

  • DALL-E 2为每个空间推理提示至少生成了一幅正确图像,表明其对物体位置与空间关系具有功能性理解。
  • 尽管如此,DALL-E 2生成的大多数图像仍为错误,表明即使在正确识别物体的情况下,也常出现对空间构型的误解。
  • ChatGPT在冯·诺依曼-摩根斯特恩理性公理上的遵循表现不一致,在许多决策场景中至少违反了一条公理。
  • 该模型的推理具有不可预测性:在简单问题上做出非理性决策,却在更复杂的赌局结构中得出了正确结论。
  • ChatGPT的输出表现出一种倾向:即使在得出正确答案时,也常使用错误的推理过程,表明其内部逻辑缺乏一致性。
  • 本研究凸显了由于大型语言模型与扩散模型本质上具有生成性和开放性特征,导致在扩展认知评估方面面临重大挑战。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。