[论文解读] Using Students as Experimental Subjects in Software Engineering Research -- A Review and Discussion of the Evidence
本文认为,在软件工程研究中使用学生作为实验对象并非本质上无效,只要适当地 contextualized,这种方法可以非常有效。基于对现有研究的综述,本文得出结论:学生参与者通常适合用于控制实验,尤其是在早期研究阶段,因为其表现与技能水平的相关性高于与职业身份的相关性,且方法论的严谨性比受试者的人口统计特征更为重要。
Should students be used as experimental subjects in software engineering? Given that students are in many cases readily available and cheap it is no surprise that the vast majority of controlled experiments in software engineering use them. But they can be argued to constitute a convenience sample that may not represent the target population (typically "real" developers), especially in terms of experience and proficiency. This causes many researchers (and reviewers) to have reservations about the external validity of student-based experiments, and claim that students should not be used. Based on an extensive review of published works that have compared students to professionals, we find that picking on "students" is counterproductive for two main reasons. First, classifying experimental subjects by their status is merely a proxy for more important and meaningful classifications, such as classifying them according to their abilities, and effort should be invested in defining and using these more meaningful classifications. Second, in many cases using students is perfectly reasonable, and student subjects can be used to obtain reliable results and further the research goals. In particular, this appears to be the case when the study involves basic programming and comprehension skills, when tools or methodologies that do not require an extensive learning curve are being compared, and in the initial formative stages of large industrial research initiatives -- in other words, in many of the cases that are suitable for controlled experiments of limited scope.
研究动机与目标
- 评估在软件工程研究中使用学生作为实验受试者的有效性与外部有效性。
- 挑战一种普遍假设,即学生本质上不能代表专业开发者。
- 主张专业知识和技能水平比职业身份更具意义的分类标准。
- 提供基于证据的指导,说明在何种情况下以及如何有效使用学生受试者进行控制实验。
- 倡导改进实验设计的方法论,而非过度关注受试者的人口统计特征。
提出的方法
- 对发表的研究进行了全面回顾,比较了学生与专业软件开发人员在控制实验中的表现。
- 分析了受试者分类(学生 vs. 专业人员)与基于技能的分类对实验结果的影响。
- 评估了在试点研究中使用学生以优化实验程序,从而为后续引入专业人员提供依据。
- 回顾了软件工程实验中的统计实践,强调了稳健的数据表示和非参数方法的必要性。
- 提出了一套框架,优先考虑复制实验和方法论多样性,而非范围有限的一次性研究。
- 强调了使用真实世界环境和工业界参与以增强外部有效性的关键作用。
实验结果
研究问题
- RQ1在什么条件下适合在软件工程研究中使用学生作为实验受试者?
- RQ2在涉及编程任务和工具评估的控制实验中,学生的表现与专业人员相比如何?
- RQ3与技能水平相比,职业身份(学生 vs. 专业人员)在多大程度上影响实验结果的外部有效性?
- RQ4哪些方法论改进可以增强使用学生受试者的实验的可靠性和可推广性?
- RQ5如何通过学生试点研究来指导并合理化针对专业开发者的大规模实验?
主要发现
- 1993年至2002年间,87%的软件工程实验使用了学生作为受试者,表明尽管存在对外部有效性的担忧,该做法仍被广泛接受。
- 学生通常是真实开发者的合理替代,尤其是在涉及基础编程或低学习曲线工具的早期研究中。
- 学生与专业人员之间存在显著的技能水平重叠,且专业知识并不必然与职业身份相关。
- 在试点研究中使用学生有助于优化实验程序,并为后续阶段引入专业人员的更高成本提供正当理由。
- 主要关注点不应是受试者身份,而应是方法论严谨性,包括适当的统计分析和数据表示。
- 在多样化环境、工具和受试者群体中进行复制实验,比单纯关注受试者人口统计特征,对建立可靠知识更为关键。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。