[Paper Review] Copied citations create renowned papers?
This paper proposes that the citation distribution in scientific literature—where a few papers receive vastly more citations than others—can be explained by a simple probabilistic model based on citation copying, not scientific merit. The model shows that when researchers randomly select papers and copy references from them (including 25% of cited references), the resulting citation patterns match empirical data, suggesting that popularity arises from stochastic copying, not genius.
Recently we discovered (cond-mat/0212043) that the majority of scientific citations are copied from the lists of references used in other papers. Here we show that a model, in which a scientist picks three random papers, cites them,and also copies a quarter of their references accounts quantitatively for empirically observed citation distribution. Simple mathematical probability, not genius, can explain why some papers are cited a lot more than the other.
Motivation & Objective
- To investigate whether citation copying, rather than scientific quality, explains the heavy-tailed distribution of citations in scientific literature.
- To model the mechanism by which scientists propagate citations by copying reference lists from other papers.
- To test whether a stochastic model based on random selection and copying can reproduce empirically observed citation distributions.
- To determine the extent to which citation popularity is driven by network effects and copying behavior rather than original impact.
- To challenge the assumption that highly cited papers are cited because they are inherently more important or innovative.
Proposed method
- Propose a stochastic model in which a scientist selects three random papers and cites them.
- Incorporate a 25% probability of copying references from each cited paper into the new reference list.
- Simulate citation propagation across a network of papers using this copying mechanism.
- Compare the resulting citation distribution from the model to real-world citation data from scientific literature.
- Use mathematical probability theory to derive the expected citation frequency distribution under the copying model.
- Validate the model by fitting it to empirical citation data and assessing goodness of fit.
Experimental results
Research questions
- RQ1To what extent does citation copying explain the heavy-tailed distribution of citations in scientific literature?
- RQ2Can a simple probabilistic model based on random selection and copying replicate observed citation patterns?
- RQ3What fraction of references must be copied to reproduce real citation distributions?
- RQ4Does the model explain why a small number of papers receive a disproportionately large number of citations?
- RQ5Is the observed citation popularity in science primarily driven by stochastic copying rather than intrinsic scientific value?
Key findings
- The model, which includes copying 25% of references from cited papers, quantitatively reproduces the empirically observed citation distribution in scientific literature.
- The citation distribution generated by the copying model closely matches real-world data, indicating that stochastic copying alone can explain citation inequality.
- The model demonstrates that high citation counts can emerge purely from random selection and copying, without requiring scientific excellence.
- The results suggest that a significant portion of citation impact is due to network effects and reference list propagation, not original contribution.
- The model's success implies that citation counts may be poor indicators of scientific quality or originality.
- The study challenges the assumption that highly cited papers are cited because they are more important, showing instead that copying behavior can create 'renowned' papers through chance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.