Skip to main content
QUICK REVIEW

[论文解读] Unsupervised embedding of trajectories captures the latent structure of scientific migration

Dakota Murray, Jisung Yoon|arXiv (Cornell University)|Dec 4, 2020
Human Mobility and Location-Based Analysis被引用 4
一句话总结

本文提出使用word2vec模型学习科学迁移轨迹的密集连续向量表示,将地理位置序列视为'句子',以捕捉文化、语言和声望等潜在结构关系。该方法在数学上等价于移动性的引力模型,并证明嵌入向量编码了多维度的迁移模式,能够实现准确的功能距离估计,并揭示机构间层次化的声望与迁移趋势。

ABSTRACT

Human migration and mobility drives major societal phenomena including epidemics, economies, innovation, and the diffusion of ideas. Although human mobility and migration have been heavily constrained by geographic distance throughout the history, advances and globalization are making other factors such as language and culture increasingly more important. Advances in neural embedding models, originally designed for natural language, provide an opportunity to tame this complexity and open new avenues for the study of migration. Here, we demonstrate the ability of the model word2vec to encode nuanced relationships between discrete locations from migration trajectories, producing an accurate, dense, continuous, and meaningful vector-space representation. The resulting representation provides a functional distance between locations, as well as a digital double that can be distributed, re-used, and itself interrogated to understand the many dimensions of migration. We show that the unique power of word2vec to encode migration patterns stems from its mathematical equivalence with the gravity model of mobility. Focusing on the case of scientific migration, we apply word2vec to a database of three million migration trajectories of scientists derived from the affiliations listed on their publication records. Using techniques that leverage its semantic structure, we demonstrate that embeddings can learn the rich structure that underpins scientific migration, such as cultural, linguistic, and prestige relationships at multiple levels of granularity. Our results provide a theoretical foundation and methodological framework for using neural embeddings to represent and understand migration both within and beyond science.

研究动机与目标

  • 开发一种方法,以捕捉超越地理距离的科学迁移的多维结构。
  • 证明迁移轨迹的word2vec嵌入能够以高保真度建模机构间的功能距离。
  • 从大规模迁移数据中揭示潜在关系,如文化亲和力、语言相似性以及机构声望。
  • 通过验证嵌入空间作为支持下游迁移模式分析与机构排名的'数字孪生',验证其有效性。
  • 在理论层面建立word2vec与人类移动中引力模型之间的联系。

提出的方法

  • 将出版记录中机构隶属关系的序列视为自然语言语料中的'句子'。
  • 应用word2vec的Skip-Gram带负采样(SGNS)变体,学习机构的密集向量表示。
  • 利用所得的嵌入向量,基于余弦相似度计算机构间的功能距离。
  • 利用嵌入空间的语义结构分析文化、语言和声望关系。
  • 通过将模型预测的功能距离与移动性引力模型的预测结果进行比较,验证模型。
  • 应用降维与聚类技术,揭示机构迁移中的层次性模式。

实验结果

研究问题

  • RQ1word2vec对迁移轨迹的嵌入能否准确捕捉机构间超越地理邻近性的功能距离?
  • RQ2所学习的嵌入在多大程度上反映了机构间的文化、语言和声望关系?
  • RQ3在预测迁移模式方面,嵌入空间与移动性引力模型相比表现如何?
  • RQ4嵌入向量的范数能否预测机构规模、声望或资金水平?
  • RQ5当在不同机构层级上分析嵌入空间时,迁移的结构模式如何显现?

主要发现

  • word2vec模型在数学上等价于移动性引力模型,为其在迁移分析中的应用提供了理论基础。
  • 基于嵌入的功能距离与真实迁移流量显著相关,捕捉了超越地理因素的多维度关系。
  • 嵌入排名与机构声望(Times排名)的Spearman等级相关系数在区域性学院中为ρ = 0.49,在研究机构中为ρ = 0.58,在政府组织中为ρ = 0.36(所有p < 0.001)。
  • 机构嵌入向量的L2范数与机构规模、科研资金(S&E)、授予的博士学位数量以及排名均显著相关(所有p < 0.001)。
  • 在30个国家中观察到组织规模与嵌入向量范数之间的凹函数关系,表明嵌入空间中机构影响力的非线性缩放。
  • 迁移模式显示,精英机构和最不具优势的机构内部流动过度代表,而中等层级机构则表现出微弱的方向性亲和力,表明存在两极化的迁移结构。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。