Skip to main content
QUICK REVIEW

[论文解读] Accelerating science with human versus alien artificial intelligences

Jamshid Sourati, James A. Evans|arXiv (Cornell University)|Apr 12, 2021
Machine Learning in Materials Science参考文献 6被引用 4
一句话总结

本文提出了一种新颖的AI框架,通过在出版物、作者和科学概念的超图中建模人类专业知识的分布——特别是科学家可能做出的认知推断——来提升科学发现效率。通过在这些专家感知的推断上训练自监督模型,该方法在药物重定位和疫苗开发中的预测精度最高提升了260%,同时还能识别出‘异类’假说——即科学上有前景但被人类忽视的假说——从而加速并实现科学进步的跃迁式突破。

ABSTRACT

Data-driven artificial intelligence models fed with published scientific findings have been used to create powerful prediction engines for scientific and technological advance, such as the discovery of novel materials with desired properties and the targeted invention of new therapies and vaccines. These AI approaches typically ignore the distribution of human prediction engines -- scientists and inventor -- who continuously alter the landscape of discovery and invention. As a result, AI hypotheses are designed to substitute for human experts, failing to complement them for punctuated collective advance. Here we show that incorporating the distribution of human expertise into self-supervised models by training on inferences cognitively available to experts dramatically improves AI prediction of future human discoveries and inventions. Including expert-awareness into models that propose (a) valuable energy-relevant materials increases the precision of materials predictions by ~100%, (b) repurposing thousands of drugs to treat new diseases increases precision by 43%, and (c) COVID-19 vaccine candidates examined in clinical trials by 260%. These models succeed by predicting human predictions and the scientists who will make them. By tuning AI to avoid the crowd, however, it generates scientifically promising "alien" hypotheses unlikely to be imagined or pursued without intervention, not only accelerating but punctuating scientific advance. By identifying and correcting for collective human bias, these models also suggest opportunities to improve human prediction by reformulating science education for discovery.

研究动机与目标

  • 解决现有AI模型在科学发现中忽略专家分布及其认知推断的局限性。
  • 通过将科学专业知识的动态分布融入自监督学习模型,提升未来科学发现的预测准确性。
  • 通过规避人类认知偏见,识别出科学上有前景但被人类忽视的假说——即‘异类’假说。
  • 证明专家感知的AI不仅能加速已知的发现路径,还能通过提出高潜力但被忽视的研究方向,实现科学进步的跃迁式突破。
  • 通过识别集体人类在科学推理中的偏见,为科学教育改革提供依据,说明如何通过AI增强的发现来纠正这些偏见。

提出的方法

  • 该方法构建了一个混合超图,其中节点代表材料、属性和研究人员,边连接在出版物中共同出现的节点集合。
  • 通过在超图上进行随机游走,建模认知上可获得的推断分布——即具有特定研究背景的科学家可能做出的推理路径。
  • 采用基于GraphSAGE的图自编码器实现自监督学习框架,学习保留超图中结构关系与专家感知关系的低维节点嵌入。
  • 模型通过链接预测损失进行训练,采用负采样策略:正样本来自随机游走的滑动窗口,负样本来自幂次为3/4的单项分布。
  • 在两种设置下进行评估:一种包含作者节点(捕捉专家感知的推理路径),另一种不包含(基线,忽略专业知识分布)。
  • 通过将预测发现与材料科学、药物重定位和疫苗开发中的已知未来发现进行比较,衡量预测精度。

实验结果

研究问题

  • RQ1将人类专业知识的分布纳入AI模型,能在多大程度上提升对未来科学发现的预测准确性?
  • RQ2AI模型在多大程度上能识别出因认知偏见或推理路径受限而不太可能被人类专家提出或追求的科学上有前景的假说?
  • RQ3专家感知的AI模型是否不仅能加速已知的发现轨迹,还能通过识别高潜力但被忽视的研究方向,实现科学进步的跃迁式突破?
  • RQ4在模型的归纳偏置中包含作者和研究历史数据,相较于忽略人类专业知识的模型,其预测未来发现的能力有何提升?
  • RQ5‘异类’AI假说——即避开人类推理模式的假说——对改革科学教育和增强集体科学创新能力有何启示?

主要发现

  • 与忽略人类专业知识的模型相比,将人类专业知识分布融入自监督模型,可使能源相关材料的预测精度提高约100%。
  • 在药物重定位任务中,专家感知模型相比未考虑人类推理模式的基线模型,预测精度提升了43%。
  • 在预测新冠疫苗临床试验候选药物时,专家感知模型相比忽略人类专业知识的模型,预测精度提升了260%。
  • 该模型成功识别出‘异类’假说——即科学上有前景但极不可能被人类想象或追求的假说——从而实现非渐进式的科学突破。
  • 通过基于研究人员研究历史建模其可获得的认知推断,该方法捕捉到一种稳定的社交事实,从而改善对哪些科学思想已被尝试并放弃的推断。
  • 结果表明,纠正科学推理中的集体人类偏见,可为科学教育改革提供依据,以增强以发现为导向的思维方式。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。