[论文解读] Applying Artificial Intelligence for Age Estimation in Digital Forensic Investigations
本文提出了一种基于CNN的年龄估计原型系统,采用一个包含327张图像(0–20岁)的新颖且多样化的面部图像数据集,旨在提升儿童性虐待与剥削案件的数字取证调查效率。该系统基于DEX算法进行训练,在青少年群体(10–20岁)中实现了1.79的平均绝对误差(MAE),在法定成年年龄边界附近表现出色,但对10岁以下儿童的年龄估计准确性仍有限。
The precise age estimation of child sexual abuse and exploitation (CSAE) victims is one of the most significant digital forensic challenges. Investigators often need to determine the age of victims by looking at images and interpreting the sexual development stages and other human characteristics. The main priority - safeguarding children -- is often negatively impacted by a huge forensic backlog, cognitive bias and the immense psychological stress that this work can entail. This paper evaluates existing facial image datasets and proposes a new dataset tailored to the needs of similar digital forensic research contributions. This small, diverse dataset of 0 to 20-year-old individuals contains 245 images and is merged with 82 unique images from the FG-NET dataset, thus achieving a total of 327 images with high image diversity and low age range density. The new dataset is tested on the Deep EXpectation (DEX) algorithm pre-trained on the IMDB-WIKI dataset. The overall results for young adolescents aged 10 to 15 and older adolescents/adults aged 16 to 20 are very encouraging -- achieving MAEs as low as 1.79, but also suggest that the accuracy for children aged 0 to 10 needs further work. In order to determine the efficacy of the prototype, valuable input of four digital forensic experts, including two forensic investigators, has been taken into account to improve age estimation results. Further research is required to extend datasets both concerning image density and the equal distribution of factors such as gender and racial diversity.
研究动机与目标
- 解决儿童性虐待与剥削(CSAE)调查中年龄估计的关键挑战,因人为判断易受认知偏见和心理压力影响。
- 开发一个专为数字取证需求设计的新颖且多样的面部图像数据集,弥补现有数据集中图像多样性不足及0–10岁年龄段分布不佳的缺陷。
- 评估预训练的Deep EXpectation(DEX)模型在新数据集上的性能,以检验其在取证场景中的准确性和可靠性。
- 整合数字 forensic 专家的反馈,优化原型系统的可用性,并确保其与真实世界调查工作流程保持一致。
- 推进AI辅助决策支持系统的发展,以减少人为接触创伤性证据的暴露,提升CSAE案件中年龄估计的速度与一致性。
提出的方法
- 从原始来源精选245张图像,并与FG-NET数据集中的82张图像合并,形成共327张图像的数据集,具有高度多样性与均衡的年龄分布。
- 采用原始在IMDB-WIKI数据集上预训练的Deep EXpectation(DEX)卷积神经网络(CNN)模型,用于在新数据集上进行年龄估计。
- 应用DEX模型从面部图像预测年龄,输出以年龄概率分布形式呈现,以支持取证决策。
- 通过平均绝对误差(MAE)进行定量评估,衡量不同年龄组的预测准确性,尤其聚焦于10–15岁和16–20岁人群。
- 通过访谈和软件评估,从四位数字 forensic 专家(包括两名调查员)处收集定性反馈,以评估系统的可用性并识别关键关切点。
- 通过整合专家意见,改进原型设计与评估框架,以应对图像质量、遮挡和种族偏见等局限性。
实验结果
研究问题
- RQ1微调后的DEX模型是否能在专为数字取证调查设计的新颖且多样的面部图像数据集上实现可靠的年龄估计?
- RQ2该年龄估计模型在不同年龄组中的表现如何,特别是在法定成年年龄(16–20岁)附近的性能表现如何?
- RQ3数字 forensic 专家在多大程度上认为AI辅助年龄估计具有价值?他们对准确性、偏见和可用性方面提出了哪些关切?
- RQ4现有面部图像数据集和预训练模型在真实世界数字 forensic 应用中存在哪些主要局限性?
- RQ5如何有效整合专家反馈以提升AI原型在 forensic 工作流中的实际应用价值?
主要发现
- DEX模型在10–20岁年龄段实现了1.79的平均绝对误差(MAE),表明其在接近法定成年年龄的青少年群体中表现优异。
- 对较年长青少年(16–20岁)的估计性能尤为突出,结果与当前最先进解决方案相当,表明其在法律成年年龄分类中具有高度可靠性。
- 0–10岁儿童的年龄估计准确性仍然较低,表明该年龄段仍需进一步优化模型并扩展数据集。
- 数字 forensic 专家对AI辅助年龄估计表示强烈支持,认为其可减轻心理负担并提升判断一致性,但对种族偏见、图像模糊和面部遮挡表示担忧。
- 由原始数据与FG-NET数据合并而成的327张图像新数据集,相比现有数据集具有更高的多样性与更低的年龄范围密度,尽管其规模仍显有限。
- 专家反馈强调,将年龄预测结果以概率分布形式可视化,有助于支持取证决策,并减少对单一数值估计的过度自信。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。