Skip to main content
QUICK REVIEW

[Paper Review] Benford's Law anomalies in the 2009 Iranian presidential election

Boudewijn F. Roukema|arXiv (Cornell University)|Jun 16, 2009
Benford’s Law and Fraud Detection19 citations
TL;DR

This paper applies Benford's Law to analyze first-digit vote frequency anomalies in the 2009 Iranian presidential election. It identifies a statistically significant excess of votes for candidate K starting with digit 7 (41 observed vs. 21.2–22 expected), which correlates with suspicious patterns in high-population districts, consistent with potential vote manipulation.

ABSTRACT

The vote count first digit frequencies of the 2009 Iranian presidential election are analysed assuming proportionality of candidates' votes to the total vote per voting area. This method is closely related to Benford's Law. A highly significant (p ~ 0.0007) excess of vote counts for candidate K that start with the digit 7 is found (41 observed, 21.2--22 expected). Using this property as a selection criterion leads to the following coincidences. (i) Among the six most populous voting areas, this criterion selects those three that have greater proportions of votes for A than the other three. The probability that the two sub-groups are drawn from the same distribution is p ~ 0.1. (ii) K's vote counts for these same three voting areas all have the same second digit. The probability of this is p ~ 0.01. (iii) Most (75%) of the vote counts for K in voting areas with 70 to 79 votes for K are odd, and every even number occurs exactly once. The probability of the latter is p ~ 0.0005. Interpreting the big city effect (i)+(ii) as an overestimate of the true vote, assumed to be roughly 50% to match other data, while retaining constant total vote numbers and increasing votes for the other three candidates in proportion to their average voting percentages, would imply that the difference between A's and M's vote totals would drop by about one million votes. These results do not exclude other anomalies.

Motivation & Objective

  • To investigate potential irregularities in the 2009 Iranian presidential election using statistical analysis based on Benford's Law.
  • To assess whether observed vote count distributions deviate significantly from expected Benford-distributed patterns.
  • To identify systematic anomalies in vote counts that may indicate manipulation or reporting bias.
  • To evaluate the consistency of observed patterns across high-population voting areas and their implications for vote totals.

Proposed method

  • Analyzes first-digit frequencies of vote counts for candidate K across voting areas, assuming proportionality to total votes per area.
  • Applies Benford's Law as a benchmark for expected digit distribution in naturally occurring data.
  • Uses statistical hypothesis testing to evaluate the significance of deviations, particularly the overrepresentation of digit 7.
  • Applies the digit-7 anomaly as a selection criterion to identify subsets of voting areas with suspicious patterns.
  • Performs comparative analysis between subgroups of high-population areas to test distributional consistency.
  • Simulates vote redistribution under a model of overestimation to assess impact on vote margins between candidates.

Experimental results

Research questions

  • RQ1Is there a statistically significant deviation from Benford's Law in the first-digit distribution of vote counts for candidate K in the 2009 Iranian election?
  • RQ2Do voting areas with high vote counts for K exhibit consistent anomalies in second digits and parity of vote counts?
  • RQ3Are the observed patterns in high-population districts consistent with random variation or indicative of systematic manipulation?
  • RQ4What would be the impact on the vote margin between candidates A and M if the anomalies were corrected under a plausible overestimation model?

Key findings

  • A highly significant excess of 41 vote counts for candidate K starting with digit 7 was observed, far exceeding the expected 21.2–22 counts under Benford's Law (p ~ 0.0007).
  • Among the six most populous voting areas, the three selected by the digit-7 criterion had a higher proportion of votes for candidate A than the other three, with a p-value of ~0.1 for this distributional difference.
  • In the three digit-7-selected areas, all of K's vote counts shared the same second digit, an event with a p-value of ~0.01.
  • In voting areas with 70–79 votes for K, 75% of the vote counts were odd, and every even number occurred exactly once, a pattern with a p-value of ~0.0005.
  • Adjusting for overestimation of K's votes in big cities—while preserving total vote counts and proportionally increasing votes for other candidates—would reduce the A–M vote margin by approximately one million votes.
  • The observed anomalies are inconsistent with random variation and suggest potential manipulation, though other irregularities may also be present.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.