[Paper Review] The Five Factor Model of personality and evaluation of drug consumption risk
This study applies the Five Factor Model (FFM) of personality, along with impulsivity and sensation-seeking traits, to predict individual risk of consuming 18 psychoactive drugs using machine learning. Using online survey data and cross-validated classification models, it achieves high accuracy—sensitivity and specificity above 75%—for cannabis, crack, ecstasy, legal highs, LSD, and volatile substance abuse.
The problem of evaluating an individual's risk of drug consumption and misuse is highly important. An online survey methodology was employed to collect data including Big Five personality traits (NEO-FFI-R), impulsivity (BIS-11), sensation seeking (ImpSS), and demographic information. The data set contained information on the consumption of 18 central nervous system psychoactive drugs. Correlation analysis demonstrated the existence of groups of drugs with strongly correlated consumption patterns. Three correlation pleiades were identified, named by the central drug in the pleiade: ecstasy, heroin, and benzodiazepines pleiades. An exhaustive search was performed to select the most effective subset of input features and data mining methods to classify users and non-users for each drug and pleiad. A number of classification methods were employed (decision tree, random forest, $k$-nearest neighbors, linear discriminant analysis, Gaussian mixture, probability density function estimation, logistic regression and na{ï}ve Bayes) and the most effective classifier was selected for each drug. The quality of classification was surprisingly high with sensitivity and specificity (evaluated by leave-one-out cross-validation) being greater than 70\% for almost all classification tasks. The best results with sensitivity and specificity being greater than 75\% were achieved for cannabis, crack, ecstasy, legal highs, LSD, and volatile substance abuse (VSA).
Motivation & Objective
- To assess the predictive power of the Five Factor Model (FFM) and related traits (impulsivity, sensation seeking) in identifying individual risk for drug consumption.
- To identify clusters of drugs with correlated consumption patterns using correlation analysis.
- To develop and evaluate machine learning models that classify individuals as users or non-users of specific drugs based on personality and demographic data.
- To determine the most effective combination of input features and classification algorithms for predicting drug use risk with high sensitivity and specificity.
- To evaluate model performance using rigorous leave-one-out cross-validation to ensure robustness and generalizability.
Proposed method
- Collected online survey data on 18 central nervous system psychoactive drugs, including NEO-FFI-R scores for the Big Five personality traits, BIS-11 for impulsivity, and ImpSS for sensation seeking.
- Performed correlation analysis to identify groups of drugs with strongly correlated consumption patterns, revealing three distinct 'pleiades': ecstasy, heroin, and benzodiazepines.
- Conducted an exhaustive search to identify the optimal subset of input features (personality traits, demographics) for each drug and pleiad.
- Evaluated eight classification methods: decision tree, random forest, k-nearest neighbors, linear discriminant analysis, Gaussian mixture, probability density function estimation, logistic regression, and naive Bayes.
- Selected the best-performing classifier for each drug and pleiad using leave-one-out cross-validation to minimize overfitting.
- Reported model performance using sensitivity and specificity as primary metrics to assess classification accuracy.
Experimental results
Research questions
- RQ1Which personality traits from the Five Factor Model are most strongly associated with the risk of consuming specific psychoactive drugs?
- RQ2Do clusters of drugs with correlated consumption patterns exist, and can they be identified through statistical correlation analysis?
- RQ3Which combination of input features (personality traits, demographics) and machine learning algorithms yields the highest classification accuracy in distinguishing drug users from non-users?
- RQ4Can machine learning models trained on personality and demographic data achieve high sensitivity and specificity in predicting drug use risk, even with limited data?
- RQ5How does model performance vary across different drugs and drug clusters (pleiades), particularly for substances like cannabis, ecstasy, and volatile substances?
Key findings
- The study identified three distinct correlation pleiades of drugs: ecstasy, heroin, and benzodiazepines, each centered on a core substance with highly correlated use patterns.
- For cannabis, crack, ecstasy, legal highs, LSD, and volatile substance abuse (VSA), the best-performing models achieved both sensitivity and specificity above 75%.
- Overall, classification performance was strong, with sensitivity and specificity exceeding 70% for nearly all drugs and pleiades, indicating high predictive power of personality traits.
- The most effective models were selected through an exhaustive search across multiple algorithms and feature subsets, with logistic regression and random forest frequently performing well.
- The results demonstrate that personality traits from the Five Factor Model, combined with impulsivity and sensation-seeking scores, are robust predictors of drug use risk.
- Leave-one-out cross-validation confirmed high model robustness, with consistent performance across different drugs and user groups.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.