[Paper Review] The effect of dataset size and the process of big data mining for investigating solar-thermal desalination by using machine learning
This study develops an optimized big data mining pipeline for solar-thermal desalination using machine learning, reducing data collection time by 83.3% via ultra-hydrophilic condensation covers and generating over 1,000 high-quality datasets. It demonstrates that Random Forest outperforms other models for datasets larger than 1,000, achieving a minimum mean relative prediction error of 4% on extrapolation, and reveals that dataset range significantly affects feature importance, with up to 115% variation in weighted values.
Machine learning's application in solar-thermal desalination is limited by data shortage and inconsistent analysis. This study develops an optimized dataset collection and analysis process for the representative solar still. By ultra-hydrophilic treatment on the condensation cover, the dataset collection process reduces the collection time by 83.3%. Over 1,000 datasets are collected, which is nearly one order of magnitude larger than up-to-date works. Then, a new interdisciplinary process flow is proposed. Some meaningful results are obtained that were not addressed by previous studies. It is found that Radom Forest might be a better choice for datasets larger than 1,000 due to both high accuracy and fast speed. Besides, the dataset range affects the quantified importance (weighted value) of factors significantly, with up to a 115% increment. Moreover, the results show that machine learning has a high accuracy on the extrapolation prediction of productivity, where the minimum mean relative prediction error is just around 4%. The results of this work not only show the necessity of the dataset characteristics' effect but also provide a standard process for studying solar-thermal desalination by machine learning, which would pave the way for interdisciplinary study.
Motivation & Objective
- To address data scarcity and inconsistent analysis in machine learning applications for solar-thermal desalination.
- To reduce the time required for dataset collection in experimental solar still studies.
- To develop a standardized, interdisciplinary big data mining process for solar-thermal desalination research.
- To investigate the impact of dataset size and range on machine learning model performance and feature importance.
- To enable accurate extrapolation predictions of desalination productivity using trained models.
Proposed method
- Applying ultra-hydrophilic treatment to the condensation cover to enhance condensation efficiency and accelerate data collection.
- Collecting over 1,000 experimental datasets—nearly tenfold more than prior studies—through an optimized data acquisition protocol.
- Implementing a novel interdisciplinary process flow integrating experimental design, data preprocessing, and machine learning model evaluation.
- Using Random Forest, XGBoost, and other models to evaluate predictive accuracy and computational speed across varying dataset sizes.
- Analyzing feature importance using permutation-based methods to assess the impact of dataset range on weighted feature contributions.
- Validating model performance on extrapolation tasks using independent test sets with low mean relative prediction error.
Experimental results
Research questions
- RQ1How does increasing dataset size affect the performance and selection of machine learning models in solar-thermal desalination prediction?
- RQ2To what extent does the range of the dataset influence the quantified importance of input factors in machine learning models?
- RQ3Can machine learning models achieve high accuracy in extrapolating productivity beyond the training data range?
- RQ4What is the optimal machine learning model for large-scale solar-thermal desalination datasets in terms of both accuracy and speed?
- RQ5How can the data collection process be optimized to significantly reduce time while maintaining data quality?
Key findings
- The ultra-hydrophilic treatment reduced data collection time by 83.3% compared to conventional methods.
- Over 1,000 high-quality datasets were collected, representing a nearly tenfold increase over existing studies.
- Random Forest emerged as the optimal model for datasets larger than 1,000, offering both high accuracy and fast inference speed.
- The range of the dataset influenced feature importance values by up to 115%, indicating strong sensitivity to data distribution.
- Machine learning models achieved a minimum mean relative prediction error of just 4% on extrapolation tasks, demonstrating strong generalization capability.
- The study establishes a standardized, reproducible process flow for machine learning applications in solar-thermal desalination research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.