Skip to main content
QUICK REVIEW

[Paper Review] Innovation with and without patents

Josef Taalbi|arXiv (Cornell University)|Oct 8, 2022
Innovation Policy and R&D5 citations
TL;DR

This study analyzes the overlap between patented inventions and actual commercialized innovations using a matched dataset of 4,460 Swedish innovations (1970–2015) linked to 13,561 patents. It finds that only 17% of innovation information is captured by patent data—implying an 83% information loss—highlighting the limitations of relying solely on patents to monitor innovation, even with quality adjustments.

ABSTRACT

A long-standing discussion is to what extent patents can be used to monitor trends in innovation activity. This study quantifies the amount and quality of information about actual innovation contained in the patent system, based on 4,460 Swedish innovations (1970-2015) that have been matched to international patents. The results show that most innovations were not patented and that among those that were, 43.9% of all innovations, only a fraction can be identified with patent quality data. The best-performing models identify 17% of all information about innovations, equivalent to an information loss of at least 83%. Econometric tests also show that the fraction of innovations responding to strengthened patent laws during the period were on average 8% percent. The overlap between the patent and innovation systems is hence more modest than often assumed. This accentuates the need to, alongside patents, develop versatile approaches in order to induce and monitor various aspects of innovation.

Motivation & Objective

  • To quantify the extent to which the patent system captures actual commercialized innovations, addressing a key gap in innovation measurement.
  • To assess the information loss in using patents as proxies for innovation activity, especially given the strategic and non-commercial nature of many patents.
  • To evaluate the effectiveness of patent quality indicators (e.g., citations) in identifying true innovations within the patent system.
  • To challenge the assumption that patents reliably reflect innovation trends, particularly in policy and economic analysis.
  • To advocate for complementary, multi-source approaches to innovation monitoring beyond patent-based metrics.

Proposed method

  • Constructed a matched dataset of 4,460 commercialized innovations from the LBIO database linked to 13,561 international patents via manual and machine-learning-assisted searches in Google Patents.
  • Applied a noisy-channel model to frame innovation-patent linkage as information transmission, with three key factors: patent propensity (ρ), recall (α), and precision (β), yielding information capture as ρ × α × β.
  • Used machine learning (Random Forest and Multilayer Perceptron) in three iterative rounds to predict innovation-patent matches, refining features over time.
  • Incorporated textual features from patent titles, abstracts, and descriptions (e.g., keyword counts, shares, vectorization), and named entity features (inventors, contacts) from innovation records.
  • Calibrated model performance using accuracy, F1-score, and error rates (false positives and negatives), with validation on manually verified pairs.
  • Conducted principal component analysis on patent features (e.g., originality, radicalness, citations) to identify structural patterns in high-impact patents.

Experimental results

Research questions

  • RQ1What fraction of actual commercialized innovations in Sweden between 1970 and 2015 are captured by the patent system?
  • RQ2To what extent do patent quality indicators (e.g., citations, originality) improve the identification of true innovations within the patent data?
  • RQ3How does the overlap between innovation and patent systems vary across sectors, time periods, and patent offices?
  • RQ4What is the information loss in relying on patents as a proxy for innovation, and how does it vary with different quality-adjusted selection methods?
  • RQ5How effective are machine learning models in identifying genuine innovation-patent linkages, and what features drive their predictive performance?

Key findings

  • Only 17% of all information about commercialized innovations is captured by the patent system, indicating a minimum information loss of 83%.
  • Among all innovations, 43.9% were patented, but only a fraction of these patents contained quality data (e.g., citations), limiting their usefulness for innovation monitoring.
  • The best-performing machine learning models achieved an F1-score of 0.795 and identified 11,694 potential innovation-patent matches, but still missed a significant share of true links.
  • Patent propensity (fraction of innovations filed as patents) was highest in the US and EPO, with a peak around 1990–2000, but remained below 50% across all sectors.
  • Patents linked to innovations received significantly more citations (mean 3.96) than non-linked patents (mean 1.93), suggesting some predictive power of citation data.
  • Principal component analysis revealed that originality, radicalness, and renewal were key features distinguishing patents linked to high-impact innovations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.