Skip to main content
QUICK REVIEW

[Paper Review] Community Targeted Phishing: A Middle Ground Between Massive and Spear Phishing through Natural Language Generation

Alberto Giaretta, Nicola Dragoni|arXiv (Cornell University)|Aug 24, 2017
Spam and Phishing Detection8 references4 citations
TL;DR

This paper introduces Community Targeted Phishing (CTP), a novel phishing approach that uses Natural Language Generation (NLG) to create machine-tailored emails for entire communities—bridging the gap between mass phishing and high-effort spear phishing. By leveraging template-driven and advanced NLG techniques, attackers can generate credible, personalized messages at scale, increasing deception effectiveness while reducing manual effort.

ABSTRACT

Looking at today phishing panorama, we are able to identify two diametrically opposed approaches. On the one hand, massive phishing targets as many people as possible with generic and preformed texts. On the other hand, spear phishing targets high-value victims with hand-crafted emails. While nowadays these two worlds partially intersect, we envision a future where Natural Language Generation (NLG) techniques will enable attackers to target populous communities with machine-tailored emails. In this paper, we introduce what we call Community Targeted Phishing (CTP), alongside with some workflows that exhibit how NLG techniques can craft such emails. Furthermore, we show how Advanced NLG techniques could provide phishers new powerful tools to bring up to the surface new information from complex data-sets, and use such information to threaten victims' private data.

Motivation & Objective

  • To propose a new phishing paradigm—Community Targeted Phishing (CTP)—that combines scalability with personalization using NLG.
  • To demonstrate how template-driven NLG can generate convincing phishing emails tailored to community-specific roles, such as researchers or colleagues.
  • To explore the potential of advanced NLG techniques in extracting and summarizing hidden data (e.g., h-index trends) to enhance phishing credibility.
  • To raise awareness about emerging threats where automated, context-aware phishing campaigns could bypass traditional spam filters.
  • To lay the foundation for future research on NLG-based phishing detection and mitigation.

Proposed method

  • Uses template-driven NLG to generate phishing emails by filling predefined templates with community-specific data (e.g., names, topics, fake links).
  • Employs workflows where templates are dynamically populated using data from sources like Google Scholar or shared document platforms.
  • Applies advanced NLG to summarize complex, heterogeneous data (e.g., academic performance trends) into natural language for social engineering.
  • Designs phishing emails that mimic trusted communication contexts, such as peer review or event invitations, to increase believability.
  • Creates fake login pages (e.g., resembling Google or social media) to harvest credentials after user interaction.
  • Leverages natural language coherence and contextual relevance to reduce suspicion and improve success rates.

Experimental results

Research questions

  • RQ1How can NLG techniques be used to generate phishing emails that are both scalable and contextually personalized for specific communities?
  • RQ2To what extent can template-driven NLG produce credible phishing messages that mimic legitimate internal communications?
  • RQ3Can advanced NLG extract and summarize hidden behavioral or performance data (e.g., h-index trends) to enhance phishing credibility?
  • RQ4How might such NLG-powered phishing campaigns bypass traditional spam filters designed for generic or overtly malicious content?
  • RQ5What are the implications of automating social engineering at scale using NLG for high-value communities like academic or professional networks?

Key findings

  • Community Targeted Phishing (CTP) enables attackers to generate large volumes of highly personalized phishing emails using NLG, targeting entire communities with minimal manual effort.
  • Template-driven NLG can produce convincing phishing messages by filling context-specific templates with data such as colleague names, research topics, and fake file-sharing links.
  • Advanced NLG techniques can summarize complex data—such as academic performance trends—into natural language, making phishing emails more credible and harder to detect.
  • Phishing emails that reference personal achievements (e.g., rising h-index) or shared interests (e.g., music genres) significantly increase the likelihood of user engagement and credential theft.
  • The integration of NLG with data mining from public sources (e.g., Google Scholar) enables automated, stealthy attacks that are more effective than generic mass phishing.
  • Traditional Bayesian spam filters may struggle to detect these attacks due to the natural language quality and contextual relevance of the generated content.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.