[论文解读] Theme and Topic: How Qualitative Research and Topic Modeling Can Be Brought Together
本文提出了 Theme and Topic 系统,通过将熟悉的定性研究概念(如编码和主题识别)映射到机器学习过程,实现了概率主题模型与定性研究工作流程的整合。该系统使研究人员能够在基于定性研究方法论的直观界面中使用自动化主题模型,提升了对算法输出的可及性与批判性参与度,同时突出了人类解释与统计建模之间的关键差异。
Qualitative research is an approach to understanding social phenomenon based around human interpretation of data, particularly text. Probabilistic topic modelling is a machine learning approach that is also based around the analysis of text and often is used to in order to understand social phenomena. Both of these approaches aim to extract important themes or topics in a textual corpus and therefore we may see them as analogous to each other. However there are also considerable differences in how the two approaches function. One is a highly human interpretive process, the other is automated and statistical. In this paper we use this analogy as the basis for our Theme and Topic system, a tool for qualitative researchers to conduct textual research that integrates topic modelling into an accessible interface. This is an example of a more general approach to the design of interactive machine learning systems in which existing human professional processes can be used as the model for processes involving machine learning. This has the particular benefit of providing a familiar approach to existing professionals, that may can make machine learning seem less alien and easier to learn. Our design approach has two elements. We first investigate the steps professionals go through when performing tasks and design a workflow for Theme and Topic that integrates machine learning. We then designed interfaces for topic modelling in which familiar concepts from qualitative research are mapped onto machine learning concepts. This makes these the machine learning concepts more familiar and easier to learn for qualitative researchers.
研究动机与目标
- 通过识别主题发现中的共同目标,弥合定性研究与自动化主题模型之间的差距。
- 设计一种交互式系统,通过将机器学习概念与既定的定性工作流程对齐,使主题模型对定性研究人员更易用。
- 解决将机器学习整合到定性实践中所面临的挑战,包括算法不透明性和用户误解。
- 通过使算法局限性可见且可解释,支持对主题模型输出的批判性参与。
- 通过将机器学习建立在定性探究的认识论基础之上,促进混合方法研究中的方法论素养。
提出的方法
- 该系统将定性研究概念(如编码、主题识别和迭代分析)映射到概率主题模型的组件上。
- 其工作流程模仿了迭代的定性过程,包括数据审查、代码生成和主题优化。
- 界面显示自动生成的主题及其关联关键词,并通过视觉提示(例如置灰的代码)表明即使在删除后,这些代码仍保有影响。
- 用户可通过排除词语、编辑代码以及查看与每个主题相关联的文本片段,交互式地优化主题。
- 设计强调对算法局限性的提示,例如词语与主题的不完美匹配,以防止误解。
- 系统围绕一个概念模型构建,将机器学习操作转化为熟悉的定性研究行为。
实验结果
研究问题
- RQ1如何在不损害解释严谨性的情况下,将主题模型有意义地整合到定性研究工作流程中?
- RQ2哪些设计原则能使定性研究人员理解并批判性评估主题模型的输出?
- RQ3主题模型的假设与行为在哪些方面与定性研究中的人类解释过程相悖?
- RQ4交互式系统如何使非技术背景的定性研究人员能够理解机器学习概念?
- RQ5哪些界面模式可以帮助用户识别并响应主题模型中的算法局限性?
主要发现
- Theme and Topic 系统成功地将定性研究实践(如编码和主题构建)映射到主题模型上,使研究人员能够在熟悉的流程中使用自动化分析。
- 用户可以直观检查每个主题关联的词语,并观察像 'dumped' 这类术语如何与主题代码相关联,即使未被显式标注。
- 被删除的代码以灰色显示,表明其仍对主题模型产生影响,防止用户误以为其语义影响已被完全移除。
- 系统突出了人类解释与算法输出之间的差异,例如当词语因统计模式而非语境意义被分配到主题时。
- 通过使算法行为透明,该设计支持批判性参与,降低了对主题模型结果不加批判依赖的风险。
- 该方法表明,当基于现有专业实践和概念模型时,将主题模型整合到定性研究中是可行的。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。