Skip to main content
QUICK REVIEW

[论文解读] Fighting Authorship Linkability with Crowdsourcing

Mishari Almishari, Ekin Oğuz|arXiv (Cornell University)|May 19, 2014
Authorship Attribution and Profiling参考文献 16被引用 5
一句话总结

本文提出通过众包和机器翻译来降低社区评论平台中的作者身份可链接性。通过让陌生人使用亚马逊MTurk重写评论或通过多种语言翻译评论,风格特征得到充分改变,使现有链接性分析技术失效,同时保持可读性和语义完整性。

ABSTRACT

Massive amounts of contributed content -- including traditional literature, blogs, music, videos, reviews and tweets -- are available on the Internet today, with authors numbering in many millions. Textual information, such as product or service reviews, is an important and increasingly popular type of content that is being used as a foundation of many trendy community-based reviewing sites, such as TripAdvisor and Yelp. Some recent results have shown that, due partly to their specialized/topical nature, sets of reviews authored by the same person are readily linkable based on simple stylometric features. In practice, this means that individuals who author more than a few reviews under different accounts (whether within one site or across multiple sites) can be linked, which represents a significant loss of privacy. In this paper, we start by showing that the problem is actually worse than previously believed. We then explore ways to mitigate authorship linkability in community-based reviewing. We first attempt to harness the global power of crowdsourcing by engaging random strangers into the process of re-writing reviews. As our empirical results (obtained from Amazon Mechanical Turk) clearly demonstrate, crowdsourcing yields impressively sensible reviews that reflect sufficiently different stylometric characteristics such that prior stylometric linkability techniques become largely ineffective. We also consider using machine translation to automatically re-write reviews. Contrary to what was previously believed, our results show that translation decreases authorship linkability as the number of intermediate languages grows. Finally, we explore the combination of crowdsourcing and machine translation and report on the results.

研究动机与目标

  • 为应对同一作者在多个账户上撰写的评论所引发的风格特征可链接性带来的日益增长的隐私风险。
  • 探究众包重写是否能有效隐藏作者的写作风格特征,同时保持评论的原始语义。
  • 评估多语言机器翻译在降低可链接性方面的影响,并评估其可读性权衡。
  • 探索结合众包与翻译的混合方法,以实现隐私与实用性之间的最佳平衡。
  • 开发一种实用的插件框架,使作者能够在评论平台上自动化完成匿名化处理。

提出的方法

  • 通过亚马逊机械土耳其人(Amazon Mechanical Turk)平台,以每项任务0.12美元的名义费用,众包随机工人重写原始评论。
  • 收集重写后的评论,并通过人工评估方法评估其可读性和与原文的语义相似度。
  • 使用机器翻译工具将评论通过多种中间语言进行翻译,以增加风格上的差异性。
  • 通过重新应用风格特征分析(使用Writeprints特征子集)测量原始、重写和翻译后评论的可链接性降低程度。
  • 将众包与机器翻译相结合,以在保持低可链接性的同时提升翻译后评论的可读性。
  • 设计了一种潜在的浏览器插件,用于自动化提交任务至众包平台并获取匿名化后的评论。

实验结果

研究问题

  • RQ1众包重写在不牺牲可读性的前提下,能在多大程度上降低社区评论中的作者身份可链接性?
  • RQ2增加中间翻译语言的数量如何影响作者身份的可链接性?
  • RQ3机器翻译后的评论是否足够可读,可作为实际的匿名化工具使用?
  • RQ4众包与机器翻译相结合对可链接性和评论质量的综合影响如何?
  • RQ5此类匿名化技术在真实世界评论平台中的部署在可行性与可扩展性方面如何?

主要发现

  • 众包重写使作者身份可链接性降低至接近无效水平,链接性分析技术在重写后的评论上准确率仅达10%。
  • 人工评估确认,重写后的评论保持了较高的可读性和与原文的语义相似度。
  • 多语言翻译显著降低了可链接性,且中间语言数量越多,混淆效果越明显。
  • 机器翻译后的评论可读性低于众包重写版本,但结合众包后可读性得到改善。
  • 众包成本较低(平均每篇评论0.12美元),且延迟在典型评论审核周期内可接受。
  • 由于平台政策及无法大规模关联原始内容与重写内容,对请求者和工人的隐私风险极低。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。