[论文解读] Research Topic Flows in Co-Authorship Networks
本文提出主题流网络(TFN),一种有向加权多图,通过合作者关系和摘要数据建模作者与研究领域之间的主题专长流动。结合非负矩阵分解进行主题建模与合作者网络分析,TFN 实现了对跨主题与同主题知识流动的大规模分析,揭示了过去 60 年计算机科学与数学领域中动态的合作模式及具有影响力的科研人员。
In scientometrics, scientific collaboration is often analyzed by means of co-authorships. An aspect which is often overlooked and more difficult to quantify is the flow of expertise between authors from different research topics, which is an important part of scientific progress. With the Topic Flow Network (TFN) we propose a graph structure for the analysis of research topic flows between scientific authors and their respective research fields. Based on a multi-graph and a topic model, our proposed network structure accounts for intratopic as well as intertopic flows. Our method requires for the construction of a TFN solely a corpus of publications (i.e., author and abstract information). From this, research topics are discovered automatically through non-negative matrix factorization. The thereof derived TFN allows for the application of social network analysis techniques, such as common metrics and community detection. Most importantly, it allows for the analysis of intertopic flows on a large, macroscopic scale, i.e., between research topic, as well as on a microscopic scale, i.e., between certain sets of authors. We demonstrate the utility of TFNs by applying our method to two comprehensive corpora of altogether 20 Mio. publications spanning more than 60 years of research in the fields computer science and mathematics. Our results give evidence that TFNs are suitable, e.g., for the analysis of topical communities, the discovery of important authors in different fields, and, most notably, the analysis of intertopic flows, i.e., the transfer of topical expertise. Besides that, our method opens new directions for future research, such as the investigation of influence relationships between research fields.
研究动机与目标
- 建模研究人员与研究领域之间的主题专长流动,这是传统合作者网络分析中常被忽视的方面。
- 开发一种可扩展且可解释的方法,用于分析无需引用数据的跨主题与同主题协作流动。
- 仅使用出版物元数据(作者与摘要信息)实现对研究领域间知识转移的宏观与微观分析。
- 展示 TFN 在大规模科学语料中检测主题社区、有影响力的作者以及随时间演变的主题流动方面的实用性。
- 在科学计量学中开启新的研究方向,特别是研究不同研究领域之间的影响关系与专长转移。
提出的方法
- 构建一个多图结构(TFN),其中有向边表示作者之间的主题协作,边权重由源作者在特定主题上的相对专长决定。
- 对摘要应用非负矩阵分解(NMF)以自动发现研究主题,并估算每位作者的主题专长。
- 将边权重定义为源作者主题专长与目标作者专长的比值,确保有向流动能反映知识转移潜力。
- 聚合作者间的主题流动,以计算宏观层面(主题到主题)的跨主题流动,并随时间进行分析。
- 在生成的 TFN 上整合标准社会网络分析技术(如 PageRank、k-core 分解和社区检测),以识别有影响力的作者和主题社区。
- 使用来自 1960 年以来计算机科学与数学领域(共 2000 万篇文献,90 万名作者)的两个大规模语料库数据,数据源为 Semantic Scholar 开放研究语料库。
实验结果
研究问题
- RQ1如何在大规模合作者网络中建模并量化研究人员与研究领域之间的主题专长流动?
- RQ2主题建模与合作者数据在多大程度上可结合以揭示随时间演变的跨主题知识转移模式?
- RQ3在真实世界科学语料库中,主题流网络在数十年间表现出怎样的结构与动态特性?
- RQ4跨主题流动在不同研究领域和时间周期内的强度与方向有何差异?
- RQ5TFN 能否基于专长流动模式识别出有影响力的作者与主题社区?其结果与标准网络指标相比如何?
主要发现
- TFN 模型成功仅使用合作者关系与摘要数据,无需引用信息,即可捕捉同主题与跨主题的专长流动。
- 该方法实现了对跨越 2000 万篇文献、历时 60 余年在计算机科学与数学领域中主题流动的大规模分析。
- 主题流网络揭示了跨主题流动的动态演变,不同时间与研究领域间流动模式存在显著差异。
- 在 TFN 上进行的社区检测与 PageRank 分析识别出的有影响力作者与主题社区,与已知的科学进展相符,验证了模型的可解释性。
- TFN 结构支持标准图论分析,证明其在检测主题协作模式方面的鲁棒性与实用性。
- 该方法为研究不同研究领域之间的因果影响(如神经网络对计算机视觉的影响)开辟了新途径,通过建模随时间演变的专长转移。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。