Skip to main content
QUICK REVIEW

[论文解读] Demand Forecasting for Platelet Usage: from Univariate Time Series to Multivariate Models

Maryam Motamedi, Dawson, Jessica|arXiv (Cornell University)|Jan 6, 2021
Blood donation and transfusion practices被引用 7
一句话总结

本研究为加拿大血液服务组织的血小板使用量提出了一种多变量需求预测框架,结合临床预测因子(如患者实验室检查结果、人口统计学信息)与时间序列数据,采用ARIMA、Prophet、套索回归、随机森林和LSTM模型进行建模。结果表明,多变量模型显著提升了预测准确性,尤其是在数据有限的情况下;而单变量模型在历史数据充足时仍具有效性,为减少血制品供应链中的浪费与短缺提供了可操作的见解。

ABSTRACT

Platelet products are both expensive and have very short shelf lives. As usage rates for platelets are highly variable, the effective management of platelet demand and supply is very important yet challenging. The primary goal of this paper is to present an efficient forecasting model for platelet demand at Canadian Blood Services (CBS). To accomplish this goal, four different demand forecasting methods, ARIMA (Auto Regressive Moving Average), Prophet, lasso regression (least absolute shrinkage and selection operator) and LSTM (Long Short-Term Memory) networks are utilized and evaluated. We use a large clinical dataset for a centralized blood distribution centre for four hospitals in Hamilton, Ontario, spanning from 2010 to 2018 and consisting of daily platelet transfusions along with information such as the product specifications, the recipients' characteristics, and the recipients' laboratory test results. This study is the first to utilize different methods from statistical time series models to data-driven regression and a machine learning technique for platelet transfusion using clinical predictors and with different amounts of data. We find that the multivariate approaches have the highest accuracy in general, however, if sufficient data are available, a simpler time series approach such as ARIMA appears to be sufficient. We also comment on the approach to choose clinical indicators (inputs) for the multivariate models.

研究动机与目标

  • 解决血制品供应链中血小板需求高度波动且保质期短(3–5天)的挑战。
  • 减少医院血制品供应系统中目前高达9–15%的浪费率以及占总量14%的紧急当日订单。
  • 通过引入历史需求以外的临床预测因子,提升需求预测的准确性。
  • 评估单变量(ARIMA、Prophet)与多变量(套索回归、随机森林、LSTM)模型在真实临床数据上的性能表现。
  • 开发一种透明、数据驱动的预测系统,以支持医院和血制品供应商更好地进行库存与资源规划。

提出的方法

  • 利用2010–2018年间来自四个汉密尔顿医院的61,377例血小板输注的大型临床数据集。
  • 应用仅依赖历史需求模式的单变量时间序列模型(ARIMA、Prophet)。
  • 实施使用临床预测因子(如实验室检查结果、患者人口统计学信息及医院特定因素)的多变量模型(套索回归、随机森林、LSTM)。
  • 采用滚动窗口交叉验证方法,评估模型在不同时间段内的性能表现。
  • 通过套索回归进行特征选择,以提升模型可解释性及深度学习模型的输入质量。
  • 使用MAE、RMSE和MAPE指标比较模型性能,以评估预测准确性。

实验结果

研究问题

  • RQ1包含临床预测因子的多变量模型与仅依赖历史需求的单变量时间序列模型在预测血小板需求方面表现如何比较?
  • RQ2数据量对单变量与多变量预测模型性能的影响是什么?
  • RQ3哪些临床预测因子显著影响血小板需求,以及如何有效选择并整合到预测模型中?
  • RQ4与单变量方法相比,多变量模型在需求高峰期间能多大程度上降低预测误差?
  • RQ5临床数据的整合在多大程度上能提升血制品供应商与医院之间在供应链中的透明度与协调性?

主要发现

  • 在数据有限的情况下,多变量模型(套索回归、随机森林、LSTM)始终优于单变量模型(ARIMA、Prophet),尤其在捕捉复杂需求模式方面表现更优。
  • 当历史数据充足时,ARIMA等单变量模型的预测准确率可与多变量模型相当甚至更高,表明其具备良好的泛化能力与效率。
  • 引入临床预测因子显著提升了预测准确性,能够捕捉非线性依赖关系及患者状况、医院位置等情境因素。
  • 所有模型在预测需求高峰时均存在低估现象,表明需要通过事后调整或基于优化的订货策略加以改进。
  • 套索回归在特征选择方面表现优异,提升了深度学习模型的输入质量,凸显其在LSTM等复杂模型预处理中的价值。
  • 本研究证明,将临床数据整合到预测模型中可减少供应链中的低效现象,包括浪费与紧急订单,从而提升需求透明度与规划水平。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。