[论文解读] Regional Development Classification Model using Decision Tree Approach
本研究提出一种基于决策树的模型,利用中爪哇省和万丹省的GDP数据对区域发展水平进行分类。采用J48算法,准确率达85.18%,识别出欠发达地区——万鸦老市、坦格朗县、肯德尔、玛杰朗、佩马拉旺、雷姆邦、三宝垄和沃诺索博,为政策制定者提供一种快速、数据驱动的决策支持工具。
Regional development classification is one way to look at differences in levels of development outcomes. Some frequently used methods are the shift share, Gain index, the Iindex Williamson and Klassen typology. The development of science in the field of data mining, offers a new way for regional development data classification. This study discusses how the decision tree is used to classify the level of development based on indicators of regional gross domestic product (GDP). GDP Data Central Java and Banten used in this study. Before the data is entered into the decision tree forming algorithm, both the provincial GDP data are classified using Klassen typology. Three decision tree algorithms, namely J48, NBTRee and REPTree tested in this study using cross-validation evaluation, then selected one of the best performing algorithms. The results show that the J48 has a better accuracy rate which is equal to 85.18% compared to the algorithm NBTRee and REPTree. Testing the model is done to the six districts / municipalities in the province of Banten, and shows that there are two districts / cities are still at the development of the status quadrant relatively underdeveloped regions, namely Kota Tangerang and Kabupaten Tangerang. As for the Central Java Province, Kendal, Magelang, Pemalang, Rembang, Semarang and Wonosobo are an area with a quadrant of development also on the status of the region is relatively underdeveloped. Classification model that has been developed is able to classify the level of development fast and easy to enter data directly into the decision tree is formed. This study can be used as an alternative decision support for policy makers in order to determine the future direction of development.
研究动机与目标
- 开发一种基于机器学习的区域发展水平数据驱动分类模型。
- 评估多种决策树算法在基于GDP指标对区域发展进行分类时的性能表现。
- 利用标准化分类框架识别中爪哇省和万丹省的欠发达地区。
- 为区域发展规划提供一种实用、快速且可扩展的决策支持系统。
- 比较J48、NBTRee和REPTree算法在区域数据分类准确率方面的表现。
提出的方法
- 本研究以中爪哇省和万丹省15个地区/市的GDP数据作为输入特征。
- 应用Klassen类型学作为初步分类方法,以结构化地区的发展状况。
- 使用10折交叉验证训练并评估三种决策树算法——J48、NBTRee和REPTree。
- 通过分类准确率衡量模型性能,并选择表现最佳的算法进行最终分析。
- 将最终模型在万丹省六个地区/市进行测试,以验证其预测能力。
- 决策树结构可生成可解释的分类规则,为政策相关洞察提供支持。
实验结果
研究问题
- RQ1基于GDP数据,哪种决策树算法在分类区域发展水平方面表现最佳?
- RQ2机器学习模型能以多高的准确率利用省级GDP指标将地区分类为发展象限?
- RQ3本模型将中爪哇省和万丹省的哪些地区/市分类为相对欠发达地区?
- RQ4所提出的模型能否作为可扩展且可解释的决策支持工具,用于区域发展政策?
- RQ5基于Klassen类型学的预处理步骤如何影响分类结果?
主要发现
- J48算法达到最高分类准确率85.18%,优于NBTRee和REPTree。
- 万丹省的万鸦老市和坦格朗县被分类为相对欠发达地区。
- 在中爪哇省,肯德尔、玛杰朗、佩马拉旺、雷姆邦、三宝垄和沃诺索博也被识别为欠发达地区。
- 该模型以高可解释性和低计算成本成功分类了区域发展水平。
- 将Klassen类型学用作预处理步骤,提高了发展状况分类的一致性。
- 最终的决策树模型适用于直接输入数据,支持快速生成政策相关的分类结果。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。