Skip to main content
QUICK REVIEW

[论文解读] Science and Ethnicity: How Ethnicities Shape the Evolution of Computer Science Research Community

Zhaohui Wu, Dayu Yuan|arXiv (Cornell University)|Nov 5, 2014
Web visibility and informetrics参考文献 26被引用 6
一句话总结

本研究采用基于姓名的族裔分类方法,分析1936年至2010年计算机科学研究群体的演变,揭示了亚裔族裔(尤其是华裔和印裔姓名)的显著增长,到2010年已占出版物近50%。研究证明,族裔是合作者网络中的同质性因素,影响合作模式与群体形成。

ABSTRACT

Globalization and the world wide web has resulted in academia and science being an international and multicultural community forged by researchers and scientists with different ethnicities. How ethnicity shapes the evolution of membership, status and interactions of the scientific community, however, is not well understood. This is due to the difficulty of ethnicity identification at the large scale. We use name ethnicity classification as an indicator of ethnicity. Based on automatic name ethnicity classification of 1.7+ million authors gathered from Web, the name ethnicity of computer science scholars is investigated by population size, publication contribution and collaboration strength. By showing the evolution of name ethnicity from 1936 to 2010, we discover that ethnicity diversity has increased significantly over time and that different research communities in certain publication venues have different ethnicity compositions. We notice a clear rise in the number of Asian name ethnicities in papers. Their fraction of publication contribution increases from approximately 10% to near 50% from 1970 to 2010. We also find that name ethnicity acts as a homophily factor on coauthor networks, shaping the formation of coauthorship as well as evolution of research communities.

研究动机与目标

  • 调查族裔如何影响计算机科学中科学群体的演变,特别是在成员构成、地位及合作方面。
  • 通过姓名族裔作为族裔身份的代理指标,弥补科学领域中大规模族裔研究的不足。
  • 分析计算机科学研究中不同族裔群体的人口动态、出版贡献及合作强度。
  • 探讨族裔是否在合作者网络中起到同质性因素的作用,以及对研究群体形成的影响。
  • 为科学政策、移民及教育资助提供关于多样性趋势及其影响的见解。

提出的方法

  • 在215,672个具有已知族裔背景的维基百科姓名上训练一个12类多项式逻辑回归分类器,以从姓名预测族裔。
  • 将该分类器应用于Arnetminer和DBLP中的170万条作者姓名,推断出版记录中的族裔身份。
  • 使用人口规模、出版产出及合作强度作为指标,追踪1936年至2010年间各族裔群体的演变。
  • 构建合作者网络,并通过基于姓名族裔配对的边密度和聚类分析来测量同质性。
  • 使用时间序列图展示人口与贡献趋势,并通过族裔分类的合作者网络聚类分析可视化。
  • 通过基准数据集验证分类器平均准确率达85%,将姓名族裔视为更广泛族裔身份的代理。

实验结果

研究问题

  • RQ11936年至2010年间,不同族裔群体在计算机科学领域的人口规模与出版贡献如何演变?
  • RQ2族裔在合作者网络中在多大程度上作为同质性因素发挥作用?
  • RQ3计算机科学内部的不同研究群体中是否存在显著不同的族裔构成?
  • RQ4亚裔姓名(尤其是华裔与印裔)的上升如何影响计算机科学整体的多样性与合作模式?
  • RQ5基于姓名的族裔分类能否揭示国家或国家层面分析所无法察觉的科学合作与群体形成中的系统性趋势?

主要发现

  • 具有亚裔姓名的作者所占出版物比例从1970年代的约10%上升至2010年的近50%,表明代表性显著提升。
  • 1990年代后,华裔与印裔姓名族裔的出版贡献显著增长,超过欧洲姓名族裔。
  • 合作者网络表现出强烈的同质性,研究人员更可能与同一名族裔者合作,可视化中表现为节点群集。
  • 计算机科学内部的不同研究群体展现出不同的族裔构成,表明各子领域在多样性与融合程度上存在差异。
  • 亚裔族裔(尤其是华裔与印裔)之间的合作强度随时间显著增强,反映出其在该领域中日益增长的融合与影响力。
  • 本研究揭示,通过姓名推断的族裔是塑造科学合作与群体演变的重要因素,为超越国籍分析提供了新见解。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。