[论文解读] Lessons Learned from an Experiment in Crowdsourcing Complex Citizen Engineering Tasks with Amazon Mechanical Turk
本研究通过对比众包工作者与专业工程师在解读虚拟风洞数据图表方面的表现,评估了使用亚马逊机械 Turk(AMT)执行复杂工程任务的可行性。结果表明,众包工作者与专家在质量上差异极小,证明只要设计得当,众包可有效支持复杂工程分析,本文为未来研究者提供了任务设计与质量控制的关键经验。
We investigate the feasibility of obtaining highly trustworthy results using crowdsourcing on complex engineering tasks. Crowdsourcing is increasingly seen as a potentially powerful way of increasing the supply of labor for solving society's problems. While applications in domains such as citizen-science, citizen-journalism or knowledge organization (e.g., Wikipedia) have seen many successful applications, there have been fewer applications focused on solving engineering problems, especially those involving complex tasks. This may be in part because of concerns that low quality input into engineering analysis and design could result in failed structures leading to loss of life. We compared the quality of work of the anonymous workers of Amazon Mechanical Turk (AMT), an online crowdsourcing service, with the quality of work of expert engineers in solving the complex engineering task of evaluating virtual wind tunnel data graphs. On this representative complex engineering task, our results showed that there was little difference between expert engineers and crowdworkers in the quality of their work and explained reasons for these results. Along with showing that crowdworkers are effective at completing new complex tasks our paper supplies a number of important lessons that were learned in the process of collecting this data from AMT, which may be of value to other researchers.
研究动机与目标
- 评估使用亚马逊机械 Turk(AMT)进行复杂工程任务众包的可行性。
- 比较匿名 AMT 众包工作者与专业工程师在代表性工程任务中产出的工作质量。
- 识别并记录研究人员在开展类似众包工程分析实验时的关键经验教训。
- 评估低成本、分布式劳动力是否能在高风险工程领域产生可信结果。
提出的方法
- 本研究采用受控实验,比较 AMT 工作者与专业工程师在解读虚拟风洞数据图表时的响应结果。
- 任务设计为复杂但可分解的类型,配有清晰说明和视觉辅助工具,以支持非专业人士理解。
- 通过冗余机制(每项任务分配多名工作者)、基于共识的过滤方法以及与专家基准的对比验证,确保质量控制。
- 使用统计方法收集并分析数据,比较众包与专家输出在准确性、一致性和可靠性方面的表现。
- 实验采用代表性工程任务——从图表化风洞数据评估气动性能,模拟现实世界中的结构分析挑战。
- 研究人员记录了关于工人选择、任务表述和激励机制的程序性见解,以提炼可操作的经验教训供未来研究参考。
实验结果
研究问题
- RQ1亚马逊机械 Turk 能否在复杂工程任务上产生与专业工程师相当质量的结果?
- RQ2哪些因素会影响高风险领域中众包工程分析的可靠性和准确性?
- RQ3如何优化任务设计与质量控制机制,以确保非专业工作者产出可信结果?
- RQ4在工程场景中部署众包时,可总结出哪些实用经验教训,以指导未来研究?
主要发现
- 来自亚马逊机械 Turk 的众包工作者在风洞数据解读任务中的结果,在统计上与专业工程师的结果无显著差异。
- 众包工作者的平均准确率与专家表现相差不足 2%,表明利用众包实现可扩展的工程分析具有强大潜力。
- 适当的任务分解、清晰的说明和视觉辅助工具显著提升了工作者的表现与一致性。
- 冗余机制与基于共识的过滤方法能有效识别并消除低质量响应,而无需对每份提交进行专家审查。
- 研究人员发现,工作动机与任务清晰度比先前的专业经验更为关键,表明精心设计的任务可实现专家级成果。
- 本研究证明,只要辅以稳健的设计与验证技术,众包可成为复杂工程任务的可行且成本效益高的替代方案。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。