[Paper Review] Formulation Graphs for Mapping Structure-Composition of Battery Electrolytes to Device Performance
This paper introduces Formulation Graph Convolution Network (F-GCN), a deep learning model that maps the structure-composition of battery electrolyte formulations to device performance by combining molecular graph embeddings with physical properties (HOMO-LUMO, dipole moment) via knowledge transfer. The model achieves state-of-the-art prediction accuracy for Coulombic efficiency and specific capacity in Li/Cu and Li-I full-cell systems, demonstrating its ability to predict new electrolyte formulations with minimal experimental validation.
Advanced computational methods are being actively sought for addressing the challenges associated with discovery and development of new combinatorial material such as formulations. A widely adopted approach involves domain informed high-throughput screening of individual components that can be combined into a formulation. This manages to accelerate the discovery of new compounds for a target application but still leave the process of identifying the right 'formulation' from the shortlisted chemical space largely a laboratory experiment-driven process. We report a deep learning model, Formulation Graph Convolution Network (F-GCN), that can map structure-composition relationship of the individual components to the property of liquid formulation as whole. Multiple GCNs are assembled in parallel that featurize formulation constituents domain-intuitively on the fly. The resulting molecular descriptors are scaled based on respective constituent's molar percentage in the formulation, followed by formalizing into a combined descriptor that represents a complete formulation to an external learning architecture. The use case of proposed formulation learning model is demonstrated for battery electrolytes by training and testing it on two exemplary datasets representing electrolyte formulations vs battery performance -- one dataset is sourced from literature about Li/Cu half-cells, while the other is obtained by lab-experiments related to lithium-iodide full-cell chemistry. The model is shown to predict the performance metrics like Coulombic Efficiency (CE) and specific capacity of new electrolyte formulations with lowest reported errors. The best performing F-GCN model uses molecular descriptors derived from molecular graphs that are informed with HOMO-LUMO and electric moment properties of the molecules using a knowledge transfer technique.
Motivation & Objective
- To address the gap in accelerating the discovery of optimal electrolyte formulations beyond individual component screening.
- To develop a machine learning framework that captures the nonlinear structure-composition-property relationships in complex liquid electrolyte formulations.
- To enable data-driven prediction of battery performance metrics like Coulombic efficiency and specific capacity from molecular constituents.
- To integrate domain-specific physical properties (HOMO-LUMO, dipole) into molecular representations for improved generalization and interpretability.
- To validate the model on both literature-derived and experimentally generated datasets for real-world relevance.
Proposed method
- The F-GCN employs multiple parallel Graph Convolutional Networks (GCNs) to featurize individual electrolyte components based on their molecular graphs.
- Molecular descriptors are enhanced with quantum chemical properties—HOMO-LUMO energy levels and electric dipole moments—using a knowledge transfer technique.
- Component descriptors are scaled by their molar percentages in the formulation to reflect contribution weight in the final formulation representation.
- A combined formulation-level descriptor is constructed by aggregating scaled molecular features into a single vector representation.
- This composite descriptor is fed into an external learning architecture for regression of battery performance metrics.
- The model is trained and evaluated on two datasets: one from literature on Li/Cu half-cells and another from lab experiments on Li-I full-cell chemistry.
Experimental results
Research questions
- RQ1Can a deep learning model effectively map the structure-composition of multi-component electrolyte formulations to their electrochemical performance?
- RQ2How does incorporating physical molecular properties (HOMO-LUMO, dipole) improve prediction accuracy compared to standard molecular descriptors?
- RQ3To what extent can F-GCN generalize to unseen electrolyte formulations across different battery chemistries?
- RQ4What is the impact of molar percentage scaling on the predictive performance of formulation-level representations?
- RQ5How does F-GCN compare to existing high-throughput screening methods in identifying high-performing electrolyte formulations?
Key findings
- The F-GCN model achieved the lowest reported prediction errors for Coulombic efficiency and specific capacity among existing methods on both the Li/Cu half-cell and Li-I full-cell datasets.
- Incorporating HOMO-LUMO and electric dipole moment information via knowledge transfer significantly improved model generalization and accuracy.
- The model demonstrated strong generalization on unseen formulations, indicating robustness to chemical diversity in electrolyte components.
- The use of molar percentage scaling improved the interpretability and predictive power of the formulation-level representation.
- The model outperformed baseline approaches using only molecular graph features, confirming the value of integrating physical chemistry properties into the representation learning process.
- The framework enables rapid virtual screening of electrolyte formulations, reducing reliance on labor-intensive experimental trials.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.