Skip to main content
QUICK REVIEW

[论文解读] ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor

Steven A. Lehr, Aylin Caliskan|arXiv (Cornell University)|Jun 20, 2024
Artificial Intelligence in Healthcare and Education被引用 5
一句话总结

论文对 GPT-3.5 与 GPT-4 在四个科学角色(图书管理员、伦理学家、数据生成者、数据预测者)中的表现进行审查,在某些领域(如降低幻觉、伦理检测)有所提升,但在预测新数据方面能力有限。

ABSTRACT

How good a research scientist is ChatGPT? We systematically probed the capabilities of GPT-3.5 and GPT-4 across four central components of the scientific process: as a Research Librarian, Research Ethicist, Data Generator, and Novel Data Predictor, using psychological science as a testing field. In Study 1 (Research Librarian), unlike human researchers, GPT-3.5 and GPT-4 hallucinated, authoritatively generating fictional references 36.0% and 5.4% of the time, respectively, although GPT-4 exhibited an evolving capacity to acknowledge its fictions. In Study 2 (Research Ethicist), GPT-4 (though not GPT-3.5) proved capable of detecting violations like p-hacking in fictional research protocols, correcting 88.6% of blatantly presented issues, and 72.6% of subtly presented issues. In Study 3 (Data Generator), both models consistently replicated patterns of cultural bias previously discovered in large language corpora, indicating that ChatGPT can simulate known results, an antecedent to usefulness for both data generation and skills like hypothesis generation. Contrastingly, in Study 4 (Novel Data Predictor), neither model was successful at predicting new results absent in their training data, and neither appeared to leverage substantially new information when predicting more versus less novel outcomes. Together, these results suggest that GPT is a flawed but rapidly improving librarian, a decent research ethicist already, capable of data generation in simple domains with known characteristics but poor at predicting novel patterns of empirical data to aid future experimentation.

研究动机与目标

  • 评估 GPT-3.5 与 GPT-4 作为研究图书管理员,通过测试书目质量与幻觉率。
  • 评估 GPT-3.5 与 GPT-4 作为研究伦理学家,通过衡量对有缺陷的研究做法的检测与纠正。
  • 评估 GPT-3.5 与 GPT-4 作为数据生成者,考察偏见再现和模拟已知结果的能力。
  • 评估 GPT-3.5 与 GPT-4 作为新数据预测者,通过对未见、真实世界数据模式的预测来测试创新性与预测效度。

提出的方法

  • 研究1(图书管理员):生成1,000条参考文献(每个主题20条,覆盖25个心理学主题),并对正确性、完整性、相关性和引用数量进行评分。
  • 研究2(伦理学家):呈现18个关于有缺陷方案的情景(明显与微妙),在216次互动中评估GPT的回应在伦理/反思质量上的表现。
  • 研究3(数据生成者):让GPT估计词嵌入式的关联并在四领域的WEAT启发式评估中复制已知偏见模式。
  • 研究4(新数据预测者):以 Project Implicit 数据为基础,任务GPT预测国家层面的态度(隐性与显性),以评估新颖性与预测有效性。
  • 定量分析包括逻辑回归、Cronbach α 信度与与现实世界数据的相关分析。

实验结果

研究问题

  • RQ1GPT 是否能够在不产生幻觉的情况下可靠地编制全面且准确的书目?
  • RQ2GPT 如何检测并处理研究方案中的伦理问题和类似p 值操纵的做法?
  • RQ3GPT 在多大程度上能模拟已知数据模式(偏见、刻板印象)并生成可信的数据?
  • RQ4GPT 是否有能力预测超出其培训数据的新颖实证模式,且GPT-3.5 与GPT-4 的表现有何差异?
  • RQ5GPT 作为通用科学助手的局限性与发展轨迹是什么?

主要发现

  • GPT-3.5 对参考文献的幻觉率为36.0%;GPT-4 的幻觉率为5.4%,GPT-4 在对虚构引用的坦诚性方面有所提升(承认虚构)84.3% 的时间相较于GPT-3.5 的12.2%。
  • GPT-4 在伦理情景回应方面的表现优于GPT-3.5,平均为明显情景8.86/10、微妙情景7.26/10,而前者分别为5.39/10和4.05/10。
  • GPT 在数据生成方面能可靠地复制已知偏见模式(如WEAT类结果)并能模拟既定结果,支持其用于初步数据生成和假设生成。
  • 在新颖数据预测任务中,GPT-3.5 与 GPT-4 对高度新颖数据的预测能力有限;与现实世界结果的相关性变化且对新颖的隐性态度下降。
  • 带有数据伦理主题的提示可提升回应质量,伦理导向的提示输出质量高于非伦理提示。
  • 总体而言,GPT 是一个有缺陷但在改进中的图书管理员,一个相对可靠的伦理工具,能够进行简单领域的数据生成,但在预测新颖的实证模式方面能力较差。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。