Skip to main content
QUICK REVIEW

[论文解读] Identifying Fake Profiles in LinkedIn

Shalinda Adikari, Kaushik Dutta|arXiv (Cornell University)|Jun 2, 2020
Spam and Phishing Detection参考文献 16被引用 48
一句话总结

本文确定了用于检测假冒个人资料所需的最小公开可用的 LinkedIn 个人资料数据,并使用数据挖掘方法实现了87%的准确率和94%的真阴性率,相较于可比方法的准确率提高约14%。

ABSTRACT

As organizations increasingly rely on professionally oriented networks such as LinkedIn (the largest such social network) for building business connections, there is increasing value in having one's profile noticed within the network. As this value increases, so does the temptation to misuse the network for unethical purposes. Fake profiles have an adverse effect on the trustworthiness of the network as a whole, and can represent significant costs in time and effort in building a connection based on fake information. Unfortunately, fake profiles are difficult to identify. Approaches have been proposed for some social networks; however, these generally rely on data that are not publicly available for LinkedIn profiles. In this research, we identify the minimal set of profile data necessary for identifying fake profiles in LinkedIn, and propose an appropriate data mining approach for fake profile identification. We demonstrate that, even with limited profile data, our approach can identify fake profiles with 87% accuracy and 94% True Negative Rate, which is comparable to the results obtained based on larger data sets and more expansive profile information. Further, when compared to approaches using similar amounts and types of data, our method provides an improvement of approximately 14% accuracy.

研究动机与目标

  • 在如 LinkedIn 这样的专业网络中,激励识别假冒资料的需求。
  • 确定用于假冒资料识别的最小公开可用 LinkedIn 个人资料数据集合。
  • 开发一个能够在数据有限的情况下识别假冒资料的数据挖掘方法。
  • 将所提方法与使用类似数据量的方法进行比较,以评估性能提升。

提出的方法

  • 识别并限定用于检测的最小公开可用 LinkedIn 个人资料数据。
  • 使用所选特征应用数据挖掘方法来识别假冒个人资料。
  • 以准确率和真阴性率来评估性能,并与更大数据基线进行比较。

实验结果

研究问题

  • RQ1识别假冒个人资料所需的 LinkedIn 个人资料数据的最小子集是什么?
  • RQ2在有限的公开数据下,假冒个人资料的识别效果如何?
  • RQ3在准确率和错误阳性方面,所提出的方法与使用类似数据的其他方法相比如何?

主要发现

  • 在有限的个人资料数据下,该方法实现了87%的准确率。
  • 真阴性率达到94%。
  • 该方法在准确率方面比使用相似量和类型数据的可比方法高出约14%。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。