Skip to main content
QUICK REVIEW

[论文解读] Detecting The Corruption Of Online Questionnaires By Artificial Intelligence

Benjamin Lebrun, Sharon Temtsin|arXiv (Cornell University)|Aug 14, 2023
Artificial Intelligence in Healthcare and EducationMedicine参考文献 74被引用 3
一句话总结

本研究调查了人类和自动化系统是否能够识别在线问卷中的AI生成回答。尽管人类在识别AI撰写的内容方面达到了76%的准确率,但检测仍不可靠;自动化AI检测系统则被发现完全不可用,对使用众包平台的在线研究的数据质量构成威胁。

ABSTRACT

Online questionnaires that use crowd-sourcing platforms to recruit participants have become commonplace, due to their ease of use and low costs. Artificial Intelligence (AI) based Large Language Models (LLM) have made it easy for bad actors to automatically fill in online forms, including generating meaningful text for open-ended tasks. These technological advances threaten the data quality for studies that use online questionnaires. This study tested if text generated by an AI for the purpose of an online study can be detected by both humans and automatic AI detection systems. While humans were able to correctly identify authorship of text above chance level (76 percent accuracy), their performance was still below what would be required to ensure satisfactory data quality. Researchers currently have to rely on the disinterest of bad actors to successfully use open-ended responses as a useful tool for ensuring data quality. Automatic AI detection systems are currently completely unusable. If AIs become too prevalent in submitting responses then the costs associated with detecting fraudulent submissions will outweigh the benefits of online questionnaires. Individual attention checks will no longer be a sufficient tool to ensure good data quality. This problem can only be systematically addressed by crowd-sourcing platforms. They cannot rely on automatic AI detection systems and it is unclear how they can ensure data quality for their paying clients.

研究动机与目标

  • 评估人类评估者是否能够可靠检测在线问卷中的AI生成回答。
  • 评估自动化AI检测系统在识别大型语言模型生成的合成文本方面的表现。
  • 探讨AI生成回答对使用众包平台的在线研究中数据质量的影响。
  • 调查在AI自动化时代,开放式问卷回答是否仍可作为确保数据完整性的可靠工具。
  • 识别为确保大规模在线研究中数据可靠性所需的根本性解决方案。

提出的方法

  • 参与者被呈现来自在线问卷的开放式回答,其中一半由GPT-3.5生成,另一半由人类参与者提供。
  • 人类评估者被要求在受控的盲测研究中区分AI生成与人工撰写的内容。
  • 将自动化AI检测系统应用于对相同回答进行AI或人工生成的分类。
  • 使用标准分类指标测量人类和自动化系统在检测准确率方面的表现。
  • 研究采用混合方法,结合定性回答分析与定量检测性能评估。
  • 实验设计包含一个平衡的数据集,共100个回答(50个来自人类,50个由AI生成),以确保统计可靠性。
Figure 1: Distribution of response length from source experiment
Figure 1: Distribution of response length from source experiment

实验结果

研究问题

  • RQ1人类评估者能否可靠检测在线问卷中的AI生成回答?
  • RQ2当前自动化AI检测系统在识别大型语言模型生成的合成文本方面有多有效?
  • RQ3AI生成回答在多大程度上损害了在线研究中的数据质量?
  • RQ4在先进大语言模型的时代,开放式问卷回答是否仍可作为可靠的数据质量检查手段?
  • RQ5为确保AI自动化在线调查中的数据完整性,需要实施哪些系统性变革?

主要发现

  • 人类评估者区分AI生成与人工撰写回答的准确率达到76%,表明检测是可能的,但不足以实现稳健的数据质量控制。
  • 自动化AI检测系统被发现完全不可用,无法可靠识别合成文本。
  • 本研究证实,大型语言模型能够生成看似合理、语境恰当的回答,与人类写作极为相似,从而破坏了传统数据验证方法。
  • AI生成回答的普遍性可能使人工审查和注意力检查失效,增加在线研究中欺诈检测的成本。
  • 目前依赖人类判断和注意力检查的做法,已不足以确保使用众包平台的在线研究中的数据质量。
  • 研究结果表明,唯有平台级干预措施——如改进检测基础设施或访问控制——才能系统性地应对AI污染对在线问卷的威胁。
Figure 2: Questionnaire to measure the quality of the text.
Figure 2: Questionnaire to measure the quality of the text.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。