Skip to main content
QUICK REVIEW

[论文解读] AI for Science: An Emerging Agenda

Philipp Berens, K. Cranmer|arXiv (Cornell University)|Mar 7, 2023
Big Data and Business Intelligence被引用 8
一句话总结

本报告总结了 Dagstuhl Seminar 22382 关于科学中的机器学习,并概述了一条将数据驱动与机制建模结合的路线图,以及面向 AI 赋能科学发现的社区建设。

ABSTRACT

This report documents the programme and the outcomes of Dagstuhl Seminar 22382 "Machine Learning for Science: Bridging Data-Driven and Mechanistic Modelling". Today's scientific challenges are characterised by complexity. Interconnected natural, technological, and human systems are influenced by forces acting across time- and spatial-scales, resulting in complex interactions and emergent behaviours. Understanding these phenomena -- and leveraging scientific advances to deliver innovative solutions to improve society's health, wealth, and well-being -- requires new ways of analysing complex systems. The transformative potential of AI stems from its widespread applicability across disciplines, and will only be achieved through integration across research domains. AI for science is a rendezvous point. It brings together expertise from $\mathrm{AI}$ and application domains; combines modelling knowledge with engineering know-how; and relies on collaboration across disciplines and between humans and machines. Alongside technical advances, the next wave of progress in the field will come from building a community of machine learning researchers, domain experts, citizen scientists, and engineers working together to design and deploy effective AI tools. This report summarises the discussions from the seminar and provides a roadmap to suggest how different communities can collaborate to deliver a new wave of progress in AI and its application for scientific discovery.

研究动机与目标

  • 需要通过新的 AI 赋能方法来应对科学复杂性,并激发其研究动机。
  • 提出跨学科、跨领域的科学 AI 路线图。
  • 突出仿真、因果性和领域知识编码等核心主题。
  • 倡导社区建设、可互操作的工具包,以及软件与数据工程的最佳实践。

提出的方法

  • 汇总 Dagstuhl 研讨会的讨论,以阐明研究议程。
  • 识别推进科学 AI 的核心主题领域:仿真、因果性、以及领域知识的编码。
  • 提出可执行步骤及在科学实践中部署 AI 工具的促进环境。
  • 倡导用户友好的工具包和标准化的软件/数据工程实践。
  • 建议在 ML 研究者、领域专家和工程师之间开展跨学科合作。
Figure 1: Models along a spectrum from classical i.i.d models to strongly mechanistic differential equation models introduce aspects of causality and symmetries to create a continuum between mechanistic and data-driven worlds. Statistical or data-driven models are weakly mechanistic (i.e. they inclu
Figure 1: Models along a spectrum from classical i.i.d models to strongly mechanistic differential equation models introduce aspects of causality and symmetries to create a continuum between mechanistic and data-driven worlds. Statistical or data-driven models are weakly mechanistic (i.e. they inclu

实验结果

研究问题

  • RQ1如何设计 AI 方法以支持对复杂系统的精细仿真和数据驱动的问询?
  • RQ2如何将数据驱动模型与机制知识有效整合,以揭示科学中的因果关系?
  • RQ3在科学工作流程中确保安全、鲁棒、并且与领域对齐地部署 AI 需要哪些策略和基础设施?
  • RQ4为跨学科加速 AI 赋能科学所需的组织与社区建设行动有哪些?
  • RQ5工具包、基准测试和治理在推动科学发现的 AI 广泛采用中扮演何种角色?

主要发现

  • AI 在自然、物理、社会、医疗和工程科学中具有变革潜力,能够从多样的数据源和尺度中获得洞察。
  • 进展取决于开发将物理定律与数据驱动学习相结合的混合模型,以及缩短数据驱动与机制建模之间的差距。
  • 将 AI 融入科学实践需要领域知识编码、人与 AI 的界面,以及知识共享机制。
  • 建立一个由 ML 研究者、领域专家、公民科学家和工程师组成的社区,对设计和部署有效的 AI 工具至关重要。
  • 可执行的步骤包括创建用户友好的工具包、在软件/数据工程中采用最佳实践,以及投资跨学科人才。
  • 路线图强调跨域协作、模型可靠性的评估,以及对不确定性和社会影响的审慎考量。
Figure 2: Strategies for integrating domain insights: including information in data and including information as prior knowledge.
Figure 2: Strategies for integrating domain insights: including information in data and including information as prior knowledge.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。