[Paper Review] Dataset Artefacts are the Hidden Drivers of the Declining Disruptiveness in Science
The authors show that the reported decline in disruption of science and technology over time is driven by outliers with zero references (CD5=1); once these artefacts are excluded or properly controlled for, the decline largely disappears.
Park et al. [1] reported a decline in the disruptiveness of scientific and technological knowledge over time. Their main finding is based on the computation of CD indices, a measure of disruption in citation networks [2], across almost 45 million papers and 3.9 million patents. Due to a factual plotting mistake, database entries with zero references were omitted in the CD index distributions, hiding a large number of outliers with a maximum CD index of one, while keeping them in the analysis [1]. Our reanalysis shows that the reported decline in disruptiveness can be attributed to a relative decline of these database entries with zero references. Notably, this was not caught by the robustness checks included in the manuscript. The regression adjustment fails to control for the hidden outliers as they correspond to a discontinuity in the CD index. Proper evaluation of the Monte-Carlo simulations reveals that, because of the preservation of the hidden outliers, even random citation behaviour replicates the observed decline in disruptiveness. Finally, while these papers and patents with supposedly zero references are the hidden drivers of the reported decline, their source documents predominantly do make references, exposing them as pure dataset artefacts.
Motivation & Objective
- Reproduce Park et al.'s disruption (CD) analysis across large citation datasets (papers and patents).
- Identify whether zero-reference entries drive observed temporal declines in CD5 values.
- Evaluate robustness of Park et al.'s controls (regression and Monte Carlo simulations) to data artefacts.
- Propose correct handling of zero-reference entries to avoid artefact-driven conclusions.
Proposed method
- Define the CDt index in a temporal directed citation network to classify forward citations within a window (CDt).
- Show that zero-reference papers/patents create a discontinuity in CDt (CDt=1 when forward citations exist).
- Extend Park et al.'s regression with a zero-references dummy to control for discontinuity and assess model fit (R2).
- Reproduce Monte Carlo rewiring analyses to test whether observed declines persist under degree-preserving random networks.
- Use multiple data sources (Web of Science, PatentsView, SciSciNet) to verify artefact-driven effects.
- Provide supplementary analyses demonstrating zero-reference artefacts across data sources.

Experimental results
Research questions
- RQ1Does the observed decline in average CD5 over time persist when zero-referenced items are properly accounted for?
- RQ2Do regression controls that include a zero-references dummy adequately address the discontinuity in CD5?
- RQ3Do Monte Carlo rewiring results still mirror a decline when zero-reference artefacts are preserved or removed?
- RQ4Are zero-reference items predominantly artefacts of metadata rather than indicators of true disruption?
- RQ5Is the observed decline consistent across multiple data sources and forward citation windows?
Key findings
- Hidden outliers with CD5=1, arising from zero-reference entries, extensively contribute to the apparent decline in disruption.
- Excluding zero-reference items or properly controlling for them largely eliminates the observed temporal decline in CD5 for papers and patents.
- Including a zero-references dummy in regression models substantially improves fit (R2 from 0.10/0.15 to 0.52/0.95 for patents/papers).
- Randomly rewired networks show a similar decline when zero-reference correspondences are maintained, indicating artefacts rather than real disruption trends.
- Across data sources, the majority of CD5=1 items with zero references still contain references in their PDFs, confirming metadata errors as the source of artefacts.
- Overall, the decline in disruption over time is attributed to data quality improvements and artefacts rather than genuine scientific or technological progress.
![Figure 2: The reason why the robustness checks in Park et al. [ 1 ] failed to detect the consequences of the hidden outliers. This figure displays how the Park et al. [ 1 ] regression adjustment (models $4$ and $8$ in Supplementary Table $1$ in [ 1 ] ) fails to control for the discontinuous effect o](https://ar5iv.labs.arxiv.org/html/2402.14583/assets/x2.png)
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.