[Paper Review] Automated Inference on Criminality using Face Images.
This study presents the first automated machine learning approach to infer criminality from still face images using four classifiers (logistic regression, KNN, SVM, CNN) on a controlled dataset of 1,856 individuals, half of whom were convicted criminals. It reveals that criminal and non-criminal faces form two distinct, concentric manifolds, with criminal faces showing significantly greater facial variation than non-criminals, suggesting a higher degree of dissimilarity in facial appearance among offenders.
We study, for the first time, automated inference on criminality based solely on still face images. Via supervised machine learning, we build four classifiers (logistic regression, KNN, SVM, CNN) using facial images of 1856 real persons controlled for race, gender, age and facial expressions, nearly half of whom were convicted criminals, for discriminating between criminals and non-criminals. All four classifiers perform consistently well and produce evidence for the validity of automated face-induced inference on criminality, despite the historical controversy surrounding the topic. Also, we find some discriminating structural features for predicting criminality, such as lip curvature, eye inner corner distance, and the so-called nose-mouth angle. Above all, the most important discovery of this research is that criminal and non-criminal face images populate two quite distinctive manifolds. The variation among criminal faces is significantly greater than that of the non-criminal faces. The two manifolds consisting of criminal and non-criminal faces appear to be concentric, with the non-criminal manifold lying in the kernel with a smaller span, exhibiting a law of normality for faces of non-criminals. In other words, the faces of general law-biding public have a greater degree of resemblance compared with the faces of criminals, or criminals have a higher degree of dissimilarity in facial appearance than normal people.
Motivation & Objective
- To investigate whether facial images alone can reliably predict criminality using machine learning.
- To control for confounding variables such as race, gender, age, and facial expression in the dataset.
- To identify structural facial features that distinguish criminals from non-criminals.
- To analyze the underlying geometric distribution of facial features in criminal and non-criminal populations.
- To assess whether facial appearance patterns reflect a fundamental difference in facial variation between criminals and law-abiding individuals.
Proposed method
- Collected and curated a dataset of 1,856 facial images, balanced for race, gender, age, and expression, with approximately half from convicted criminals.
- Trained four supervised machine learning classifiers: logistic regression, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), and Convolutional Neural Network (CNN).
- Extracted and analyzed facial landmarks to identify structural features such as lip curvature, eye inner corner distance, and nose-mouth angle.
- Applied manifold learning techniques to analyze the geometric distribution of facial features in high-dimensional space.
- Compared the intrinsic dimensionality and spread of facial feature manifolds between criminal and non-criminal groups.
- Used statistical analysis to evaluate the significance of observed differences in facial variation between the two groups.
Experimental results
Research questions
- RQ1Can facial images alone be used to predict criminality with high accuracy using machine learning?
- RQ2Which facial structural features are most predictive of criminality in the absence of demographic or behavioral data?
- RQ3Do criminal faces form a distinct and more variable manifold compared to non-criminal faces in facial feature space?
- RQ4Is there a systematic difference in facial similarity between criminals and non-criminals, with non-criminals showing greater facial resemblance?
- RQ5Do the facial feature manifolds of criminals and non-criminals exhibit a concentric geometric structure?
Key findings
- All four classifiers—logistic regression, KNN, SVM, and CNN—achieved consistent and strong performance in distinguishing criminals from non-criminals based solely on facial images.
- Criminal faces exhibit significantly greater variation in facial appearance compared to non-criminal faces, indicating higher dissimilarity among offenders.
- The faces of non-criminals form a more compact, kernel-like manifold with lower intrinsic dimensionality, reflecting a 'law of normality' in facial appearance.
- Key structural features such as lip curvature, eye inner corner distance, and nose-mouth angle were found to be discriminative for criminality prediction.
- The facial feature manifolds of criminals and non-criminals were found to be concentric, with the non-criminal manifold lying at the center with a smaller span.
- The results provide empirical evidence for the existence of a distinct facial appearance pattern associated with criminality, despite historical controversy in the field.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.