Skip to main content
QUICK REVIEW

[论文解读] Heart Disease Detection using Quantum Computing and Partitioned Random Forest Methods

Hanif Heidari, Gerhard Hellstern|arXiv (Cornell University)|Aug 17, 2022
Artificial Intelligence in Healthcare被引用 5
一句话总结

本文提出一种使用2–4量子比特的混合量子随机森林(HQRF)模型,以提升早期心脏病检测的性能,通过将量子计算与分块随机森林相结合,增强准确性和鲁棒性。该模型在Cleveland数据集上达到96.43%的AUC,在Statlog数据集上达到97.78%的AUC,优于先前的混合量子神经网络(HQNN),在小样本和大样本数据集上均展现出更强的抗异常值能力与更高的效率。

ABSTRACT

Heart disease morbidity and mortality rates are increasing, which has a negative impact on public health and the global economy. Early detection of heart disease reduces the incidence of heart mortality and morbidity. Recent research has utilized quantum computing methods to predict heart disease with more than 5 qubits and are computationally intensive. Despite the higher number of qubits, earlier work reports a lower accuracy in predicting heart disease, have not considered the outlier effects, and requires more computation time and memory for heart disease prediction. To overcome these limitations, we propose hybrid random forest quantum neural network (HQRF) using a few qubits (two to four) and considered the effects of outlier in the dataset. Two open-source datasets, Cleveland and Statlog, are used in this study to apply quantum networks. The proposed algorithm has been applied on two open-source datasets and utilized two different types of testing strategies such as 10-fold cross validation and 70-30 train/test ratio. We compared the performance of our proposed methodology with our earlier algorithm called hybrid quantum neural network (HQNN) proposed in the literature for heart disease prediction. HQNN and HQRF outperform in 10-fold cross validation and 70/30 train/test split ratio, respectively. The results show that HQNN requires a large training dataset while HQRF is more appropriate for both large and small training dataset. According to the experimental results, the proposed HQRF is not sensitive to the outlier data compared to HQNN. Compared to earlier works, the proposed HQRF achieved a maximum area under the curve (AUC) of 96.43% and 97.78% in predicting heart diseases using Cleveland and Statlog datasets, respectively with HQNN. The proposed HQRF is highly efficient in detecting heart disease at an early stage and will speed up clinical diagnosis.

研究动机与目标

  • 为解决现有基于量子的心脏病预测模型存在的局限性,包括对高数量量子比特的需求、异常值处理能力差以及计算成本高等问题。
  • 开发一种更高效且鲁棒的量子机器学习模型,适用于小样本和大样本训练数据集。
  • 通过将量子计算与分块随机森林方法相结合,提升预测准确性和AUC。
  • 在保持高性能的同时降低对大规模训练数据的依赖,从而支持更早的临床诊断。
  • 评估该模型在异常值存在情况下的鲁棒性,相较于先前的混合量子神经网络模型。

提出的方法

  • 所提出的HQRF模型将量子线路与分块随机森林集成,以增强特征学习与分类性能。
  • 量子线路采用2–4个量子比特实现,最大限度降低资源消耗,同时保持高性能。
  • 数据集被划分为多个子集,每个子集由随机森林框架内的量子增强决策树处理。
  • 通过随机森林的集成特性减轻异常值的影响,降低对极端值的敏感性。
  • 采用两种评估策略:10折交叉验证与70-30训练/测试集划分,确保性能评估的稳健性。
  • 模型在两个开源数据集(Cleveland与Statlog)上进行训练与测试,涵盖多样化的临床数据分布。

实验结果

研究问题

  • RQ1是否能够通过少于5个量子比特的量子机器学习模型,在心脏病检测中实现高于现有基于量子模型的准确性?
  • RQ2在存在数据异常值的情况下,所提出的HQRF模型相较于HQNN模型表现如何?
  • RQ3HQRF模型是否在小样本和大样本训练数据集上均保持高性能?
  • RQ4采用分块随机森林架构对模型泛化能力与计算效率有何影响?
  • RQ5HQRF模型在AUC与训练效率方面相较于HQNN表现如何?

主要发现

  • HQRF模型在Cleveland数据集上实现了96.43%的AUC,优于先前的量子模型。
  • 在Statlog数据集上,HQRF实现了97.78%的AUC,表明其在跨数据集评估中表现更优。
  • HQRF对异常值数据的敏感性显著低于HQNN模型,展现出更强的鲁棒性。
  • HQRF在小样本与大样本训练数据集上均保持高性能,而HQNN仅在大规模数据集下表现最佳。
  • 该模型仅使用2–4个量子比特即实现高准确率,相比先前的量子模型显著降低了计算成本与内存占用。
  • 在10折交叉验证中,HQRF优于HQNN;而在70-30训练/测试集划分中,HQNN表现更优,表明HQNN对数据集大小存在依赖性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。