Skip to main content
QUICK REVIEW

[논문 리뷰] Intelligent Materials Modelling: Large Language Models Versus Partial Least Squares Regression for Predicting Polysulfone Membrane Mechanical Performance

Dingding Cao, Mieow Kee Chan|arXiv (Cornell University)|2026. 03. 14.
Machine Learning in Materials Science인용 수 0
한 줄 요약

본 연구는 구조적 서술자로부터 폴리설폰 멤브레인의 기계적 특성을 예측하기 위해 네 가지 LLM을 PLS와 비교 벤치마킹했고, LLM이 파단 연신률(EL)을 특히 개선하는 반면 PLS는 영 모듈러스와 인장강도 같은 선형 특성에서 여전히 경쟁력이 있음을 발견했다.

ABSTRACT

Predicting the mechanical properties of polysulfone (PSF) membranes from structural descriptors remains challenging due to extreme data scarcity typical of experimental studies. To investigate this issue, this study benchmarked knowledge-driven inference using four large language models (LLMs) (DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, and GPT-5) against partial least squares (PLS) regression for predicting Young's modulus (E), tensile strength (TS), and elongation at break (EL) based on pore diameter (PD), contact angle (CA), thickness (T), and porosity (P) measurements. These knowledge-driven approaches demonstrated property-specific advantages over the chemometric baseline. For EL, LLMs achieved statistically significant improvements, with DeepSeek-R1 and GPT-5 delivering 40.5% and 40.3% of Root Mean Square Error reductions, respectively, reducing mean absolute errors from $11.63\pm5.34$% to $5.18\pm0.17$%. Run-to-run variability was markedly compressed for LLMs ($\leq$3%) compared to PLS (up to 47%). E and TS predictions showed statistical parity between approaches ($q\geq0.05$), indicating sufficient performance of linear methods for properties with strong structure-property correlations. Error topology analysis revealed systematic regression-to-the-mean behavior dominated by data-regime effects rather than model-family limitations. These findings establish that LLMs excel for non-linear, constraint-sensitive properties under bootstrap instability, while PLS remains competitive for linear relationships requiring interpretable latent-variable decompositions. The demonstrated complementarity suggests hybrid architectures leveraging LLM-encoded knowledge within interpretable frameworks may optimise small-data materials discovery.

연구 동기 및 목표

  • 희소한 실험 데이터로부터 PSF 멤브레인의 기계적 특성을 예측하는 도전을 조사한다.
  • PD, CA, T, P를 기반으로 E(Young's 모듈러스), TS(인장강도), EL(파단 연신률)에 대해 네 가지 LLM을 이용한 지식 주도 추론을 PLS 회귀와 비교한다.
  • 부트스트랩 불안정성 하에서 모델의 신뢰성을 이해하기 위해 실행 간 변동성과 오차 특성을 평가한다.
  • 특성별 장점과 재료 발견에서 하이브리드형의 해석 가능한 아키텍처 가능성을 식별한다.

제안 방법

  • 4가지 LLM(DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, GPT-5)을 PLS 회귀와 대조 평가한다.
  • PD, CA, T, P로부터 영 모듈러스(E), 인장강도(TS), 파단 연신률(EL)을 예측한다.
  • RMSE, MAE 등의 성능 지표를 계산하고 q-값 등을 통해 통계적 유의성을 비교한다.
  • 오차 토폴로지를 분석하여 회귀-대-평균 효과와 데이터-레짐의 영향을 이해한다.

실험 결과

연구 질문

  • RQ1제한된 데이터에서 LLM이 PLS를 능가하여 PSF 멤브레인 기계적 특성 예측에 우수한가?
  • RQ2어떤 특성(E, TS, EL)이 LLM에 비해 PLS에서 가장 큰 개선을 보이는가?
  • RQ3부트스트랩 불안정성 하에서 LLM과 PLS 간 실행 간 변동성은 어떻게 비교되는가?
  • RQ4오차 토폴로지가 데이터-레짐 효과와 모델 계열의 한계에 대해 무엇을 보여주는가?
  • RQ5작은 데이터 재료 발견을 더 향상시키기 위해 LLM-활용 해석 가능한 하이브리드 프레임워크가 가능한가?

주요 결과

  • LLMs은 특성별 이점을 보이며 특히 EL에서 두드러진 성능 향상을 보이고, DeepSeek-R1에서 RMSE가 40.5%, GPT-5에서 40.3% 감소했다.
  • EL의 평균 절대 오차가 LLM 사용 시 11.63±5.34%에서 5.18±0.17%로 감소한다.
  • LLMs는 실행 간 변동성이 현저히 낮아 (≤3%) PLS의 최대 47%에 비해 작다.
  • E와 TS에 대한 예측은 LLM과 PLS 간에 통계적 동등성을 보이며(q≥0.05).
  • 오차 토폴로지는 데이터-레짐 효과에 의해 주도되는 평균으로의 회귀 행동을 나타내며, 내재된 모델 한계가 아님; LLM은 비선형이면서 제약에 민감한 특성에서 뛰어나고 PLS는 선형 관계에서 여전히 경쟁력을 유지한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.