Skip to main content
QUICK REVIEW

[论文解读] Is this word borrowed? An automatic approach to quantify the likeliness of borrowing in social media

Jasabanta Patro, Bidisha Samanta|arXiv (Cornell University)|Mar 15, 2017
Linguistics, Language Diversity, and Identity参考文献 23被引用 3
一句话总结

本文提出了一种新颖的计算框架,用于自动量化社交媒体中词汇借用的可能性,采用基于上下文的聚类方法,并引入三种新指标——UUR、UUR-adj 和 UUR-adj-2——这些指标源自英语-印地语混合语 tweets 中的用户级词汇使用模式。该方法与人工标注的真实数据相比,斯皮尔曼等级相关系数达到 0.62,超过基线水平(0.26)两倍以上,且在年轻用户和低混合使用用户中表现更优,表明能够早期检测到借用趋势。

ABSTRACT

Code-mixing or code-switching are the effortless phenomena of natural switching between two or more languages in a single conversation. Use of a foreign word in a language; however, does not necessarily mean that the speaker is code-switching because often languages borrow lexical items from other languages. If a word is borrowed, it becomes a part of the lexicon of a language; whereas, during code-switching, the speaker is aware that the conversation involves foreign words or phrases. Identifying whether a foreign word used by a bilingual speaker is due to borrowing or code-switching is a fundamental importance to theories of multilingualism, and an essential prerequisite towards the development of language and speech technologies for multilingual communities. In this paper, we present a series of novel computational methods to identify the borrowed likeliness of a word, based on the social media signals. We first propose context based clustering method to sample a set of candidate words from the social media data.Next, we propose three novel and similar metrics based on the usage of these words by the users in different tweets; these metrics were used to score and rank the candidate words indicating their borrowed likeliness. We compare these rankings with a ground truth ranking constructed through a human judgment experiment. The Spearman's rank correlation between the two rankings (nearly 0.62 for all the three metric variants) is more than double the value (0.26) of the most competitive existing baseline reported in the literature. Some other striking observations are, (i) the correlation is higher for the ground truth data elicited from the younger participants (age less than 30) than that from the older participants, and (ii )those participants who use mixed-language for tweeting the least, provide the best signals of borrowing.

研究动机与目标

  • 开发一种自动方法,用于量化语言中外来词是否为借用而非通过代码切换使用。
  • 从非正式、大规模的社交媒体数据中识别出多语言用户早期的借用信号。
  • 通过利用社交媒体中的使用模式,特别是极少进行代码混合行为的用户的数据,改进现有基线方法。
  • 通过从不同年龄群体收集的人工标注真实数据,对方法进行验证。

提出的方法

  • 采用基于上下文的聚类方法,从英语-印地语混合语 tweets 中采样 57 个候选英文词汇,确保语言相关性和多样性。
  • 提出三种新颖指标——UUR、UUR-adj 和 UUR-adj-2——基于用户层面的词汇频率和分布,用于估计借用可能性。
  • 这些指标计算用户间使用的一致性,同时对词汇频率和用户活跃度进行调整,以减少噪声和偏差。
  • 开展了一项人工评估研究,邀请 58 名评审员对 57 个目标词汇的借用可能性进行排序,按年龄组分层。
  • 使用斯皮尔曼等级相关系数将模型的排序结果与人工标注的真实数据进行比较,并进一步按用户代码混合行为进行分析。
  • 在不同用户群体(高、中、低代码混合)中评估该方法,以评估信号强度和鲁棒性。

实验结果

研究问题

  • RQ1社交媒体数据能否提供可靠的早期借用信号,以区别于代码切换?
  • RQ2社交媒体中外来词的使用模式与人工判断的借用可能性之间是否存在相关性?
  • RQ3极少进行代码混合的用户是否比频繁代码切换者提供更强的借用识别信号?
  • RQ4所提出的方法是否对年轻用户更有效,反映出借用的早期阶段趋势?
  • RQ5所提出的指标在预测借用可能性方面,与现有基线相比在定量上表现如何?

主要发现

  • 所提出的 UUR 指标与人工标注的真实数据相比,斯皮尔曼等级相关系数达到 0.62,超过基线水平(0.26)两倍以上。
  • 在年轻参与者(年龄 < 30 岁)中,相关系数显著更高,达到 0.62,表明该方法能有效捕捉早期借用趋势。
  • 当 UUR 指标基于低代码混合行为用户的数据推导时表现最佳,相关系数达 0.65,表明此类用户提供了最纯净的借用信号。
  • 精确率和召回率指标显示,UUR 相较于基线有持续改进,尤其在 SB(极少借用)和 LM(可能被借用)类别中表现突出。
  • 无论在年轻用户还是年长用户群体中,UUR 的宏平均和微平均精确率与召回率均持续高于基线水平。
  • 该方法展现出强鲁棒性和泛化能力,在多种评估方案和真实数据变体中均表现优异。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。