[论文解读] Teaching Responsible Data Science: Charting New Pedagogical Territory
本文提出了一套教学框架,通过将技术编码与伦理批判相结合,聚焦于通过“营养标签”——一组可解释性工具——实现透明度与可解释性,以教授负责任的数据科学(RDS)。该方法在纽约大学的研究生/本科生课程中进行了测试,采用基于项目的教学法,并引入“用于解释的对象”模型,产生了可操作的教学方法和公开可获取的材料,从而弥合工程与社会科学教学法之间的差距。
Although numerous ethics courses are available, with many focusing specifically on technology and computer ethics, pedagogical approaches employed in these courses rely exclusively on texts rather than on software development or data analysis. Technical students often consider these courses unimportant and a distraction from the "real" material. To develop instructional materials and methodologies that are thoughtful and engaging, we must strive for balance: between texts and coding, between critique and solution, and between cutting-edge research and practical applicability. Finding such balance is particularly difficult in the nascent field of responsible data science (RDS), where we are only starting to understand how to interface between the intrinsically different methodologies of engineering and social sciences. In this paper we recount a recent experience in developing and teaching an RDS course to graduate and advanced undergraduate students in data science. We then dive into an area that is critically important to RDS -- transparency and interpretability of machine-assisted decision-making, and tie this area to the needs of emerging RDS curricula. Recounting our own experience, and leveraging literature on pedagogical methods in data science and beyond, we propose the notion of an "object-to-interpret-with". We link this notion to "nutritional labels" -- a family of interpretability tools that are gaining popularity in RDS research and practice. With this work we aim to contribute to the nascent area of RDS education, and to inspire others in the community to come together to develop a deeper theoretical understanding of the pedagogical needs of RDS, and contribute concrete educational materials and methodologies that others can use. All course materials are publicly available at https://dataresponsibly.github.io/courses.
研究动机与目标
- 解决技术性、实践性RDS课程缺乏的问题,平衡伦理与编程,反驳技术学生认为伦理课程无关紧要的刻板印象。
- 开发一种教学框架,通过实用且基于研究的方法,将可解释性与透明度整合到数据科学教育中。
- 创建一个可扩展、可重用的课程模型,结合编程、批判性思维与实际应用,为学生做好负责任数据科学实践的准备。
- 通过建构主义和探究式学习原则,弥合工程与社会科学在RDS教育中的方法论差距。
- 提供具体、公开可获取的教育材料与评估策略,支持其他教育工作者的采用与改编。
提出的方法
- 在纽约大学开发了一门研究生/本科生RDS课程,将伦理、公平性、隐私与透明度整合到技术性数据科学训练中。
- 引入“用于解释的对象”概念——一种受“用于思考的对象”启发的建构主义学习工具——帮助学生分析和解释机器学习模型。
- 采用“营养标签”作为一组可解释性工具,用于评估模型透明度,使学生能够评估并提升模型的可解释性。
- 设计基于项目的教学活动,包括模型复现、流程分析以及可解释性标签的设计,以培养实践技能。
- 整合教学最佳实践,如示范示例、小组问题解决、模拟实验和反思,以支持多样化学习者。
- 采用混合方法评估,以衡量模型性能与可解释性,承认量化指标与定性解释之间的张力。
实验结果
研究问题
- RQ1如何在技术编程与伦理批判之间实现平衡,以提升技术学生对负责任数据科学教育的相关性?
- RQ2哪些教学模型能有效教授数据科学学生在人机决策中的可解释性与透明度?
- RQ3‘营养标签’在基于课程的学习环境中如何作为实用的可解释性工具发挥作用?
- RQ4‘用于解释的对象’在帮助学生深入理解模型透明度与公平性方面发挥什么作用?
- RQ5如何设计RDS课程,以整合技术与社会科学视角,同时保持可扩展性与可重用性?
主要发现
- 该RDS课程通过整合编程、模型复现与可解释性任务,成功吸引了技术学生,反驳了伦理课程无关紧要的刻板印象。
- 学生通过实际操作‘营养标签’和模型分析,表现出更强的识别透明度缺陷并提升可解释性能力。
- ‘用于解释的对象’框架在帮助学生从新手到专家水平理解模型可解释性方面被证明是有效的。
- 基于项目的教学,包括复现研究与可解释性标签设计,显著提升了学生参与度,并促进了公平性与透明度方面的实践技能发展。
- 评估结果揭示了基于指标的性能与定性可解释性之间的张力,凸显了在RDS教育中采用混合方法评估的必要性。
- 课程材料,包括所有作业与资源,已公开发布于 https://dataresponsibly.github.io/courses,可供教育工作者重用与改编。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。