Skip to main content
QUICK REVIEW

[论文解读] Can AI Serve as a Substitute for Human Subjects in Software Engineering Research?

Marco Aurélio Gerosa, Bianca Trinkenreich|arXiv (Cornell University)|Nov 18, 2023
Ethics and Social Impacts of AI被引用 4
一句话总结

这篇视觉论文提出使用大型语言模型(LLMs)如ChatGPT生成合成定性数据,作为软件工程研究中人类受试者的可扩展替代方案。通过采用基于角色的提示、多角色对话和超角色调查,LLMs能够模拟多样化的用户视角,从而实现对访谈、焦点小组和调查的高效数据收集——为传统方法提供了有前景的补充,但由于人类共情与细微差别的不可替代性,尚不能完全取代。

ABSTRACT

Research within sociotechnical domains, such as Software Engineering, fundamentally requires a thorough consideration of the human perspective. However, traditional qualitative data collection methods suffer from challenges related to scale, labor intensity, and the increasing difficulty of participant recruitment. This vision paper proposes a novel approach to qualitative data collection in software engineering research by harnessing the capabilities of artificial intelligence (AI), especially large language models (LLMs) like ChatGPT. We explore the potential of AI-generated synthetic text as an alternative source of qualitative data, by discussing how LLMs can replicate human responses and behaviors in research settings. We examine the application of AI in automating data collection across various methodologies, including persona-based prompting for interviews, multi-persona dialogue for focus groups, and mega-persona responses for surveys. Additionally, we discuss the prospective development of new foundation models aimed at emulating human behavior in observational studies and user evaluations. By simulating human interaction and feedback, these AI models could offer scalable and efficient means of data generation, while providing insights into human attitudes, experiences, and performance. We discuss several open problems and research opportunities to implement this vision and conclude that while AI could augment aspects of data gathering in software engineering research, it cannot replace the nuanced, empathetic understanding inherent in human subjects in some cases, and an integrated approach where both AI and human-generated data coexist will likely yield the most effective outcomes.

研究动机与目标

  • 为应对在软件工程定性研究中招募人类参与者,尤其是代表性不足群体的日益增长的挑战。
  • 探索人工智能生成的合成文本是否可作为人类来源定性数据的可行替代或补充。
  • 提出实用方法,利用基础模型模拟常见定性研究方法中的真实人类行为与反应。
  • 识别在社会技术软件工程研究中整合AI生成数据与人类生成数据时的开放问题与研究机会。
  • 倡导一种整合方法,使AI与人类数据共存,以提升研究的可扩展性与有效性。

提出的方法

  • 使用基于角色的提示,引导LLMs生成反映特定人口统计或心理画像特征的响应,例如非技术用户或资深开发人员。
  • 采用多角色对话技术,通过生成具有不同观点的多个虚拟参与者的响应,模拟焦点小组的互动。
  • 创建超角色响应,以模拟聚合用户档案的调查响应,从而实现调查研究的大规模数据生成。
  • 利用LLMs的 few-shot 及 few-shot few-shot 能力,对响应进行微调,以确保一致性、风格和行为的真实性。
  • 设计系统性变化的提示,以激发反映不同态度、动机和经验的响应,例如开源软件贡献者的动机。
  • 使用真实贡献者调查中的统计分布(例如,23.1% 的贡献者经验少于3年)来指导并支撑合成响应分布的构建。
Figure 1: Using prompt engineering in a large language model to interview a specific persona. The conversation was generated using GPT-4.
Figure 1: Using prompt engineering in a large language model to interview a specific persona. The conversation was generated using GPT-4.

实验结果

研究问题

  • RQ1LLMs能否生成准确反映真实人类受试者在软件工程研究中观点、动机与行为的合成定性数据?
  • RQ2基于角色的提示如何用于在虚拟访谈和焦点小组中模拟多样化的用户画像?
  • RQ3AI生成的响应在多大程度上能复制真实人类响应在调查和观察研究中的变异性与细微差别?
  • RQ4AI生成数据在捕捉共情、情境敏感性以及文化细微差别的真实人类体验方面存在哪些局限性?
  • RQ5如何有意义地整合AI生成与人类来源的定性数据,以提升研究的可扩展性与有效性?

主要发现

  • LLMs能够生成与人类视角高度相似的合成响应,例如估计约40%的长期开源软件贡献者可能强烈认同‘他们贡献是因为专有软件无法解决某些问题’。
  • 基于角色的提示方法可一致模拟用户类型,例如估计约50%经验少于3年的贡献者可能强烈认同‘他们贡献是为了提升技能’。
  • 多角色对话可模拟焦点小组中的群体动态,其响应反映出不同经验水平与动机的多样化观点。
  • 超角色响应可生成具有统计上合理分布的类似调查的数据,例如估计约30%的贡献者可能对某项陈述‘部分认同’,15%保持中立。
  • 该方法在保持与现实世界用户行为和动机相关性的同时,展现出大规模收集定性数据的潜力。
  • 尽管具备这些能力,研究结论仍认为AI无法完全取代人类受试者,因为合成响应缺乏真正的共情与情境深度。
Figure 2: Tweaking the persona to interviewing a woman contributor. The conversation was generated using GPT-4.
Figure 2: Tweaking the persona to interviewing a woman contributor. The conversation was generated using GPT-4.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。