[Paper Review] Intelligent Materials Modelling: Large Language Models Versus Partial Least Squares Regression for Predicting Polysulfone Membrane Mechanical Performance
The study benchmarks four LLMs against PLS for predicting polysulfone membrane mechanical properties from structural descriptors, finding LLMs especially improve elongation at break while PLS remains competitive for linear properties like Young's modulus and tensile strength.
Predicting the mechanical properties of polysulfone (PSF) membranes from structural descriptors remains challenging due to extreme data scarcity typical of experimental studies. To investigate this issue, this study benchmarked knowledge-driven inference using four large language models (LLMs) (DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, and GPT-5) against partial least squares (PLS) regression for predicting Young's modulus (E), tensile strength (TS), and elongation at break (EL) based on pore diameter (PD), contact angle (CA), thickness (T), and porosity (P) measurements. These knowledge-driven approaches demonstrated property-specific advantages over the chemometric baseline. For EL, LLMs achieved statistically significant improvements, with DeepSeek-R1 and GPT-5 delivering 40.5% and 40.3% of Root Mean Square Error reductions, respectively, reducing mean absolute errors from $11.63\pm5.34$% to $5.18\pm0.17$%. Run-to-run variability was markedly compressed for LLMs ($\leq$3%) compared to PLS (up to 47%). E and TS predictions showed statistical parity between approaches ($q\geq0.05$), indicating sufficient performance of linear methods for properties with strong structure-property correlations. Error topology analysis revealed systematic regression-to-the-mean behavior dominated by data-regime effects rather than model-family limitations. These findings establish that LLMs excel for non-linear, constraint-sensitive properties under bootstrap instability, while PLS remains competitive for linear relationships requiring interpretable latent-variable decompositions. The demonstrated complementarity suggests hybrid architectures leveraging LLM-encoded knowledge within interpretable frameworks may optimise small-data materials discovery.
Motivation & Objective
- Investigate the challenge of predicting PSF membrane mechanical properties from sparse experimental data.
- Compare knowledge-driven inference using four LLMs against PLS regression for E, TS, and EL based on PD, CA, T, and P.
- Assess run-to-run variability and error characteristics to understand model reliability under bootstrap instability.
- Identify property-specific advantages and potential for hybrid, interpretable architectures in materials discovery.
Proposed method
- Evaluate four LLMs (DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, GPT-5) against PLS regression.
- Predict Young's modulus (E), tensile strength (TS), and elongation at break (EL) from pore diameter (PD), contact angle (CA), thickness (T), and porosity (P).
- Compute performance metrics (e.g., RMSE, MAE) and compare statistical significance (e.g., via q-values).
- Analyze error topology to understand regression-to-the-mean effects and data-regime influences.
Experimental results
Research questions
- RQ1Do LLMs outperform PLS in predicting PSF membrane mechanical properties from limited data?
- RQ2Which properties (E, TS, EL) show the most improvement with LLMs versus PLS?
- RQ3How do run-to-run variabilities compare between LLMs and PLS under bootstrap instability?
- RQ4What does error topology reveal about data-regime effects versus model-family limitations?
- RQ5Can a hybrid LLM-enabled, interpretable framework further improve small-data materials discovery?
Key findings
- LLMs show property-specific advantages, notably for EL, with significant RMSE reductions for DeepSeek-R1 (40.5%) and GPT-5 (40.3%).
- Mean absolute error for EL drops from 11.63±5.34% to 5.18±0.17% with LLMs.
- LLMs exhibit markedly lower run-to-run variability (≤3%) compared to PLS (up to 47%).
- Predictions for E and TS show statistical parity between LLMs and PLS (q≥0.05).
- Error topology indicates regression-to-the-mean behavior driven by data-regime effects rather than inherent model limitations; LLMs excel for non-linear, constraint-sensitive properties, while PLS remains competitive for linear relationships.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.