[Paper Review] Integration of knowledge and data in machine learning
This paper surveys how knowledge discovery and knowledge embedding can be integrated with data-driven ML, outlines methods, gaps, and opportunities, and argues for a closed loop between discovery and embedding.
Scientific research's mandate is to comprehend and explore the world, as well as to improve it based on experience and knowledge. Knowledge embedding and knowledge discovery are two significant methods of integrating knowledge and data. Through knowledge embedding, the barriers between knowledge and data can be eliminated, and machine learning models with physical common sense can be established. Meanwhile, humans' understanding of the world is always limited, and knowledge discovery takes advantage of machine learning to extract new knowledge from observations. Knowledge discovery can not only assist researchers to better grasp the nature of physics, but it can also support them in conducting knowledge embedding research. A closed loop of knowledge generation and usage are formed by combining knowledge embedding with knowledge discovery, which can improve the robustness and accuracy of models and uncover previously unknown scientific principles. This study summarizes and analyzes extant literature, as well as identifies research gaps and future opportunities.
Motivation & Objective
- Clarify the distinction and interaction between knowledge discovery and knowledge embedding.
- Summarize existing methods for discovering governing equations from data (structure and coefficients).
- Summarize knowledge embedding techniques that integrate domain knowledge into ML models.
- Identify research gaps and opportunities for advancing integrated knowledge-data learning.
Proposed method
- Classify knowledge discovery methods into closed library, expandable library, and open-form approaches for equation mining.
- Discuss approaches for mining equations with complex structures and coefficients, including sparse regression, genetic algorithms, symbolic regression, and PDE-Net variants.
- Describe knowledge embedding strategies across data preprocessing, model structure design, and penalty/reward (soft vs hard constraints).
- Compare soft and hard constraint frameworks and their implications for data efficiency and physical fidelity.
- Highlight practical embedding techniques such as PINN, TgNN, PgNN, and physics-constrained loss formulations.
Experimental results
Research questions
- RQ1What are the main categories and capabilities of knowledge discovery methods for extracting governing equations from data?
- RQ2How can domain knowledge be embedded into ML models to improve accuracy, robustness, and physical consistency?
- RQ3What are the key challenges limiting the integration of knowledge (discovery and embedding) with data in ML, and what future directions address them?
- RQ4What is the role of coefficients (constant, expressible, inexpressible) in equation mining and how can they be effectively inferred?
- RQ5How can a closed loop between knowledge discovery and knowledge embedding be realized to advance scientific and engineering tasks?
Key findings
- Knowledge discovery methods differ by how they represent equations (closed libraries, expandable libraries, open-form) and by coefficient complexity.
- Open-form methods offer greater flexibility for complex structures but come with higher computational cost.
- Knowledge embedding can be implemented via data preprocessing, network design, and constrained optimization, with soft versus hard constraints affecting data needs and fidelity.
- Hard constraints can reduce data requirements but rely on correct domain knowledge; soft constraints are easier to implement but may not enforce physical laws exactly.
- The authors identify five gaps/opportunities for knowledge discovery (gradient-appropriate embeddings, necessary-condition extraction, complex-structure/coefficient mining, gradient accuracy, equation simplification) and five for knowledge embedding (handling complex governing equations, irregular fields via graph nets, adaptive hyperparameters, noisy/scarce data, automated ML for accessibility).
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.