[论文解读] Cluster Analysis of Educational Data: an Example of Quantitative Study on the answers to an Open-Ended Questionnaire
本文提出了一种用于教育研究中开放式问卷回答的定量分析聚类分析方法,采用无监督学习在无先验假设的情况下识别出具有智力一致性的学生群体。该方法应用于大学物理学生对问题的回答,揭示了不同的推理模式,并通过数据驱动的分类方式为学生概念理解提供了新见解。
In the last years many studies examined the consistency of students' answers in a variety of contexts. Some of these papers tried to develop more detailed models of the consistency of students' reasoning, or to subdivide a sample of students into intellectually similar subgroups. The problem of taking a set of data and separating it into subgroups where the elements of each subgroup are more similar to each other than they are to elements not in the subgroup has been extensively studied through the methods of Cluster Analysis. This method can separate students into groups that can be recognized and characterized by common traits in their answers, without any prior knowledge of what form those groups would take (unbiased classification). In this paper we start from a detailed analysis of the data coding needed in Cluster Analysis, in order to discuss the meaning and the limits of the interpretation of quantitative results. Then two methods commonly used in Cluster Analysis are described and the variables and parameters involved are outlined and criticized. Section III deals with the application of these methods to the analysis of data from an open-ended questionnaire administered to a sample of university students, and the quantitative results are discussed. Finally, the quantitative results are related to student answers and compared with previous results reported in the literature, by pointing out the new insights resulting from the application of such new methods.
研究动机与目标
- 开发一种系统方法,利用定量聚类分析对开放式问卷中的质性教育数据进行分析。
- 在不预先假设群体结构的前提下,基于学生推理模式的相似性识别出智力同质的学生子群体。
- 评估聚类分析结果在教育情境下的可解释性与局限性,特别是数据编码和参数选择方面。
- 将定量结果与现有文献进行比较,揭示聚类分析在物理教育中学生推理方面的新见解。
- 提供一种将聚类分析应用于教育数据的方法论框架,支持数据驱动的分类与有意义的解释。
提出的方法
- 本研究采用层次聚类和k-means聚类算法,基于从编码数据中提取的语义相似性对学生的回答进行分组。
- 通过详细且系统化的过程进行数据编码,将自由文本回答转化为适合聚类分析的数值变量。
- 对层次聚类中的距离度量和链接准则的选择进行批判性评估,以分析其对群体形成和可解释性的影响。
- 对k-means方法采用不同数量的聚类进行应用,并利用内部验证指数评估结果,以确定最优的群体结构。
- 在聚类前对变量进行标准化,以确保不同回答维度的贡献相等,并检验参数选择的敏感性。
- 聚类的解释基于代表性学生回答的质性分析,将数值分组与实际的推理模式相联系。
实验结果
研究问题
- RQ1如何在教育研究中有效应用聚类分析处理开放式问卷回答?
- RQ2在教育情境中,为聚类分析编码质性数据时面临的主要挑战和局限性是什么?
- RQ3从物理学生回答的聚类分析中,哪些稳定且可解释的学生推理特征浮现出来?
- RQ4所识别出的聚类与文献中已报告的学生推理类别相比如何?
- RQ5聚类分析在多大程度上揭示了超越传统质性编码的新见解,以增进对学生理解的认识?
主要发现
- 聚类分析成功识别出基于学生对开放式物理问题回答推理模式的若干明显且可解释的学生群体。
- 本研究揭示了多种推理特征,包括替代性概念和更复杂的概念框架,这些特征通过传统质性编码难以单独识别。
- 聚类参数的选择(如链接准则和聚类数量)显著影响最终的群体结构和解释结果。
- 数据编码过程被证明至关重要:编码决策的微小变化会导致不同的聚类配置,凸显了透明度和一致性的必要性。
- 该方法能够识别出以往被低估的具有相似误解或推理策略的学生子群体,为针对性教学提供了新途径。
- 聚类分析的结果与早期质性研究一致,但更为细致,提供了更系统化和可扩展的学生推理分析方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。