Skip to main content
QUICK REVIEW

[Paper Review] Accelerating science with human versus alien artificial intelligences

Jamshid Sourati, James A. Evans|arXiv (Cornell University)|Apr 12, 2021
Machine Learning in Materials Science6 references4 citations
TL;DR

This paper proposes a novel AI framework that improves scientific discovery by modeling the distribution of human expertise—specifically, the cognitive inferences scientists are likely to make—through a hypergraph of publications, authors, and scientific concepts. By training self-supervised models on these expert-aware inferences, the method boosts prediction precision by up to 260% in drug repurposing and vaccine development, while also identifying 'alien' hypotheses—scientifically promising but overlooked by humans—thereby accelerating and punctuating scientific progress.

ABSTRACT

Data-driven artificial intelligence models fed with published scientific findings have been used to create powerful prediction engines for scientific and technological advance, such as the discovery of novel materials with desired properties and the targeted invention of new therapies and vaccines. These AI approaches typically ignore the distribution of human prediction engines -- scientists and inventor -- who continuously alter the landscape of discovery and invention. As a result, AI hypotheses are designed to substitute for human experts, failing to complement them for punctuated collective advance. Here we show that incorporating the distribution of human expertise into self-supervised models by training on inferences cognitively available to experts dramatically improves AI prediction of future human discoveries and inventions. Including expert-awareness into models that propose (a) valuable energy-relevant materials increases the precision of materials predictions by ~100%, (b) repurposing thousands of drugs to treat new diseases increases precision by 43%, and (c) COVID-19 vaccine candidates examined in clinical trials by 260%. These models succeed by predicting human predictions and the scientists who will make them. By tuning AI to avoid the crowd, however, it generates scientifically promising "alien" hypotheses unlikely to be imagined or pursued without intervention, not only accelerating but punctuating scientific advance. By identifying and correcting for collective human bias, these models also suggest opportunities to improve human prediction by reformulating science education for discovery.

Motivation & Objective

  • To address the limitation of existing AI models in scientific discovery that ignore the distribution of human experts and their cognitive inferences.
  • To improve prediction accuracy of future scientific discoveries by incorporating the evolving distribution of scientific expertise into self-supervised learning models.
  • To identify scientifically promising but human-overlooked hypotheses—'alien' hypotheses—by avoiding human cognitive biases.
  • To demonstrate that expert-aware AI can both accelerate known discovery paths and punctuate progress by suggesting high-potential, overlooked research directions.
  • To inform science education reform by identifying how collective human biases in scientific reasoning can be corrected through AI-augmented discovery.

Proposed method

  • The method constructs a mixed hypergraph with nodes for materials, properties, and researchers, where edges connect sets of nodes that co-occur in publications.
  • It uses random walks over the hypergraph to model the distribution of cognitively available inferences—paths of reasoning likely to be made by scientists with specific research backgrounds.
  • A self-supervised learning framework, implemented via GraphSAGE-based graph autoencoders, learns low-dimensional node embeddings that preserve structural and expert-aware relationships in the hypergraph.
  • The model is trained using a link-prediction loss with negative sampling, where positive pairs are drawn from sliding windows over random walks and negative pairs from a unigram distribution raised to the power of 3/4.
  • The approach is evaluated in two settings: one including author nodes (capturing expert-aware inference paths) and one excluding them (baseline, ignoring expertise distribution).
  • Precision is measured by comparing predicted discoveries against known future discoveries in materials science, drug repurposing, and vaccine development.

Experimental results

Research questions

  • RQ1How does incorporating the distribution of human expertise into AI models improve prediction accuracy for future scientific discoveries?
  • RQ2To what extent can AI models identify scientifically promising hypotheses that are unlikely to be pursued by human experts due to cognitive biases or limited inference paths?
  • RQ3Can expert-aware AI models not only accelerate known discovery trajectories but also punctuate scientific progress by identifying high-potential, overlooked research avenues?
  • RQ4How does the inclusion of author and research history data in the model’s inductive bias affect its ability to predict future discoveries compared to models that ignore human expertise?
  • RQ5What are the implications of 'alien' AI hypotheses—those that avoid human inference patterns—for reforming science education and enhancing collective scientific innovation?

Key findings

  • Incorporating the distribution of human expertise into self-supervised models increases prediction precision for energy-relevant materials by approximately 100% compared to models that ignore human expertise.
  • For drug repurposing, the expert-aware model improves prediction precision by 43% over baseline models that do not account for human inference patterns.
  • In predicting clinical trial candidates for COVID-19 vaccines, the expert-aware model increases precision by 260% compared to models ignoring human expertise.
  • The model successfully identifies 'alien' hypotheses—scientifically promising but unlikely to be imagined or pursued by humans—thereby enabling punctuated rather than incremental scientific advance.
  • By modeling the cognitive inferences available to scientists based on their research history, the method captures a stable social fact that improves inference about which scientific ideas have been tried and abandoned.
  • The results suggest that correcting for collective human bias in scientific reasoning can inform reforms in science education to enhance discovery-oriented thinking.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.