Skip to main content
QUICK REVIEW

[论文解读] Machine learning for risk assessment in gender-based crime

Ángel González-Prieto, Antonio Brú|arXiv (Cornell University)|Jun 22, 2021
Crime Patterns and Interventions参考文献 38被引用 8
一句话总结

本文提出了一种用于预测亲密伴侣暴力针对妇女(IPVAW)再犯风险的机器学习(ML)模型,基于超过40,000起案件的数据集。使用欧几里得度量和收缩阈值 = 5的最近中心(NC)模型在警察保护得分上达到最高分(1.595),优于现有统计方法,并支持在执法部门中实现混合过渡模型的实用化。

ABSTRACT

Gender-based crime is one of the most concerning scourges of contemporary society. Governments worldwide have invested lots of economic and human resources to radically eliminate this threat. Despite these efforts, providing accurate predictions of the risk that a victim of gender violence has of being attacked again is still a very hard open problem. The development of new methods for issuing accurate, fair and quick predictions would allow police forces to select the most appropriate measures to prevent recidivism. In this work, we propose to apply Machine Learning (ML) techniques to create models that accurately predict the recidivism risk of a gender-violence offender. The relevance of the contribution of this work is threefold: (i) the proposed ML method outperforms the preexisting risk assessment algorithm based on classical statistical techniques, (ii) the study has been conducted through an official specific-purpose database with more than 40,000 reports of gender violence, and (iii) two new quality measures are proposed for assessing the effective police protection that a model supplies and the overload in the invested resources that it generates. Additionally, we propose a hybrid model that combines the statistical prediction methods with the ML method, permitting authorities to implement a smooth transition from the preexisting model to the ML-based model. This hybrid nature enables a decision-making process to optimally balance between the efficiency of the police system and aggressiveness of the protection measures taken.

研究动机与目标

  • 开发一种机器学习模型,以改进现有用于预测亲密伴侣暴力针对妇女(IPVAW)再犯风险的统计风险评估工具。
  • 使用反映现实世界警察保护效果和资源效率的新指标评估模型性能。
  • 提出一种混合模型,以实现从传统基于规则的系统到执法部门中基于机器学习的风险预测的平稳过渡。

提出的方法

  • 本研究使用了来自西班牙官方来源的超过40,000起IPVAW报案记录数据集,包含结构化的临床和法律变量。
  • 训练并评估了一系列机器学习模型——最近中心(NC)、K近邻(KNN)、逻辑回归(LR)、决策树(DT)和神经网络(NN)——进行比较。
  • 通过10折交叉验证对最近中心模型进行微调,以警察保护得分为优化目标函数,优化度量标准(欧几里得、曼哈顿、闵可夫斯基)和收缩阈值。
  • 通过随机插值方式构建了一种新型混合模型,将传统VioGen基于规则的系统与表现最佳的机器学习模型(NC)的预测结果结合,参数化为概率μ。
  • 该混合模型利用二项分布随机变量对两模型预测结果的差异进行扰动,确保从旧系统到新系统的平稳过渡。
  • 引入了两个新评估指标:警察保护得分(衡量有效风险检测能力)和资源过载指数(衡量过度保护措施的风险)。

实验结果

研究问题

  • RQ1机器学习模型能否在预测亲密伴侣暴力针对妇女(IPVAW)受害者再犯风险方面优于经典统计风险评估工具?
  • RQ2如何在评估模型性能时不仅关注准确率,还考虑实际操作影响,如有效保护和资源效率?
  • RQ3在不破坏现有执法实践的前提下,将新型基于机器学习的风险预测模型整合到现有执法系统中的最佳方式是什么?
  • RQ4在IPVAW再犯风险预测背景下,哪种机器学习算法在预测性能与操作可行性之间实现了最佳平衡?
  • RQ5结合传统系统与基于机器学习的预测的混合模型,能否实现安全、渐进式的过渡,同时保持或改善保护结果?

主要发现

  • 使用欧几里得度量和收缩阈值 = 5的最近中心(NC)模型在10折交叉验证中达到最高的警察保护得分为1.595,显著优于基线VioGen系统(得分为1.28)。
  • NC模型表现出强鲁棒性和一致性,在10次交叉验证折中标准差仅为0.037,表明其性能稳定。
  • 混合模型实现了从传统VioGen系统到基于机器学习模型的平稳过渡,其中μ = 0对应原始系统,μ = 1对应完整机器学习模型。
  • 所提出的警察保护得分指标有效捕捉了风险预测模型在现实世界中的实用性,倾向于正确识别高风险案例且不过度占用资源的模型。
  • 本研究引入了一种新的资源过载指数,用于量化过度保护措施的风险,有助于在安全与系统效率之间取得平衡。
  • 研究结果挑战了以往认为标准机器学习模型不适用于基于性别的犯罪风险评估的观点,证明通过适当的资料和评估方法,机器学习可显著提升预测准确性和实际操作相关性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。