Skip to main content
QUICK REVIEW

[论文解读] Onomastics 2.0 - The Power of Social Co-Occurrences

Folke Mitzlaff, Gerd Stumme|arXiv (Cornell University)|Mar 3, 2013
Names, Identity, and Discrimination Research参考文献 1被引用 4
一句话总结

本文提出一种数据挖掘方法,利用维基百科和推特的社交网络共现网络,通过基于网络的相似性度量,发现给定姓名与城市名称之间的语义关系。研究发现共现图中的结构相似性与语义相关性之间存在强烈相关性,尤其在维基百科数据中表现显著,由此催生了名为 Nameling 的成功姓名搜索平台,在四个月内吸引了超过 30,000 名用户。

ABSTRACT

Onomastics is "the science or study of the origin and forms of proper names of persons or places." ["Onomastics". Merriam-Webster.com, 2013. http://www.merriam-webster.com (11 February 2013)]. Especially personal names play an important role in daily life, as all over the world future parents are facing the task of finding a suitable given name for their child. This choice is influenced by different factors, such as the social context, language, cultural background and, in particular, personal taste. With the rise of the Social Web and its applications, users more and more interact digitally and participate in the creation of heterogeneous, distributed, collaborative data collections. These sources of data also reflect current and new naming trends as well as new emerging interrelations among names. The present work shows, how basic approaches from the field of social network analysis and information retrieval can be applied for discovering relations among names, thus extending Onomastics by data mining techniques. The considered approach starts with building co-occurrence graphs relative to data from the Social Web, respectively for given names and city names. As a main result, correlations between semantically grounded similarities among names (e.g., geographical distance for city names) and structural graph based similarities are observed. The discovered relations among given names are the foundation of "nameling" [http://nameling.net], a search engine and academic research platform for given names which attracted more than 30,000 users within four months, underpinningthe relevance of the proposed methodology.

研究动机与目标

  • 开发一种基于社交网络数据发现姓名之间新兴关系的方法论。
  • 评估基于网络的相似性度量在捕捉命名实体之间语义相关性方面的有效性。
  • 比较不同语言版本维基百科生成的共现网络,评估其结构特性和语义属性。
  • 构建一个基于社交网络分析与信息检索技术的实用姓名推荐系统。
  • 探索共现网络在按语言和文化起源对命名实体进行分类方面的潜力。

提出的方法

  • 从维基百科和推特数据构建共现图,其中节点代表姓名或城市名称,边表示在文本中共同出现。
  • 基于共享邻居,应用标准相似性函数(余弦相似性和 Jaccard 系数)来测量节点之间的结构相似性。
  • 将观察到的相似性趋势与零模型图进行比较,以排除统计伪影并评估显著性。
  • 使用最短路径距离作为网络接近度的代理指标,并将其与语义相关性(如城市之间的地理距离,姓名之间的语义相似性)进行相关性分析。
  • 通过将相似性度量与外部语义度量(如地理距离或已知姓名变体)的相关性进行比较,评估其性能。
  • 利用 Nameling 平台收集真实世界使用数据,以在推荐场景中验证相似性函数。

实验结果

研究问题

  • RQ1从维基百科和推特中提取的共现图的结构相似性在多大程度上反映了姓名之间的语义相关性?
  • RQ2基于共享邻居的相似性度量在多大程度上与城市名称的地理接近性等已知语义关系相关?
  • RQ3维基百科与推特的共现网络在捕捉的语义方面有何差异?
  • RQ4维基百科共现网络中的语言特异性特征能否用于按语言或文化起源对姓名进行分类?
  • RQ5在真实世界姓名推荐场景中,基本的网络相似性函数有多有效?

主要发现

  • 对于基于维基百科的共现网络,随着最短路径距离的增加,平均相似性(余弦相似性和 Jaccard 系数)呈强烈单调递减,表明相邻姓名在语义上更为相似。
  • 维基百科网络中的直接邻居在语义上显著优于随机节点对,且距离为二的节点对相似性低于随机预期,证实结构相似性能够捕捉语义相关性。
  • 在维基百科网络中,最短路径距离与城市名称的地理距离之间存在正相关关系,且对大多数距离而言具有统计显著性。
  • 相比之下,推特网络中城市名称的最短路径距离与地理距离呈相反关系,表明推特中共同出现的语义与维基百科存在根本性差异。
  • 基于所发现的姓名关系构建的 Nameling 平台,在四个月内吸引了超过 30,000 名用户,证明了该方法论的实际相关性。
  • 特征向量中心性分析揭示了多语言维基百科共现网络中存在明显的语言特异性特征,表明其在语言感知的实体分类方面具有潜力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。