[Paper Review] Email in the Era of LLMs
The paper presents HR Simulator™, a game to study human–LLM email writing, revealing a hybrid human+LLM advantage and how model size affects email judgments, tone, and tact.
Email communication increasingly involves large language models (LLMs), but we lack intuition on how they will read, write, and optimize for nuanced social goals. We introduce HR Simulator, a game where communication is the core mechanic: players play as a Human Resources officer and write emails to solve socially challenging workplace scenarios. An analysis of 600+ human and LLM emails with LLMs-as-judge reveals evidence for larger LLMs becoming more homogenous in their email quality judgments. Under LLM judges, humans underperform LLMs (e.g., 23.5% vs. 48-54% success rate), but a human+LLM approach can outperform LLM-only (e.g., from 40% to nearly 100% in one scenario). In cases where models' email preferences disagree, emergent tact is a plausible explanation: weaker models prefer less tactful strategies while stronger models prefer more tactful ones. Regarding tone, LLM emails are more formal and empathetic while human emails are more varied. LLM rewrites make human emails more formal and empathetic, but models still struggle to imitate human emails in the low empathy, low formality quadrant, which highlights a limitation of current post-training approaches. Our results demonstrate the efficacy of communication games as instruments to measure communication in the era of LLMs, and posit human-LLM co-writing as an effective form of communication in that future.
Motivation & Objective
- Motivate understanding of how LLMs read, write, and optimize emails to social goals in workplace contexts.
- Introduce HR Simulator™ to measure and compare human, AI, and hybrid email writing under varied scenarios.
- Characterize how LLM judges' opinions on email quality converge as models scale.
- Explore how tone, empathy, formality, and tact influence email effectiveness under AI judgment.
- Provide implications for future human–LLM collaboration in email communication.
Proposed method
- Develop HR Simulator™, a game where players act as a Human Resources officer and write emails to solve workplace scenarios.
- Use GPT-4o as the in-game judge to simulate recipients and outcomes across five scenarios.
- Analyze over 600 human and LLM emails, evaluated by multiple LLM judges from small to large models.
- Apply Elo ranking to compare judge-pair preferences for emails within the same scenario.
- Annotate emails for tact, empathy, and formality to interpret tone and alignment with model preferences.
- Perform post-hoc analyses to assess how judge size and agreement affect pass rates and perceived quality.

Experimental results
Research questions
- RQ1How do human and LLM emails compare in success rates across socially challenging workplace scenarios?
- RQ2Do larger LLMs converge toward more uniform judgments of email quality, and how does this affect preferences for AI-written content?
- RQ3Can human+LLM collaboration outperform either humans or LLMs alone in producing effective emails?
- RQ4What roles do tact, empathy, and formality play in model judgments of email quality?
- RQ5Are there systemic gaps in current post-training approaches for producing low-empathy, low-formality emails?
Key findings
- Humans alone achieve a 23.5% pass rate on average, while top LLMs reach 48–54%; human+LLM rewrites can outperform both in some scenarios.
- LLM judges rate LLM-written emails higher than human-written ones, and human+LLM emails can outperform both in certain cases.
- As model size increases, LLM judges become more homogeneous in their quality judgments, reaching about 0.5 Krippendorff’s alpha for agreement.
- Weaker judges prefer more direct emails, while stronger judges prefer more tactful, subtle emails, a phenomenon termed emergent tact.
- LLM rewrites tend to make human emails more formal and empathetic, moving them toward the high-empathy, high-formality quadrant; however, LLMs struggle to imitate low-empathy, low-formality emails.
- Human–LLM hybrid advantage arises because rewritten human emails can fall into GPT-4o’s preferred tact range, enhancing pass rates for several judges (e.g., GPT-4o and Claude 3.5 Haiku in Scenario 1).

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.