[Paper Review] Simple models predict behavior at least as well as behavioral scientists
This study evaluates whether behavioral scientists predict human behavior more accurately than simple models, finding they consistently underperform compared to basic heuristics like 'behavioral interventions have no effect' or random chance. Using data from five studies involving over 600 experts, the research shows behavioral scientists' predictions are no better than—and often worse than—simple models, revealing significant noise and overconfidence in expert judgment.
How accurately can behavioral scientists predict behavior? To answer this question, we analyzed data from five studies in which 640 professional behavioral scientists predicted the results of one or more behavioral science experiments. We compared the behavioral scientists' predictions to random chance, linear models, and simple heuristics like "behavioral interventions have no effect" and "all published psychology research is false." We find that behavioral scientists are consistently no better than - and often worse than - these simple heuristics and models. Behavioral scientists' predictions are not only noisy but also biased. They systematically overestimate how well behavioral science "works": overestimating the effectiveness of behavioral interventions, the impact of psychological phenomena like time discounting, and the replicability of published psychology research.
Motivation & Objective
- To assess the accuracy of behavioral scientists' predictions about human behavior compared to simple models and heuristics.
- To investigate whether expert judgment in behavioral science outperforms basic benchmarks such as random chance or null models.
- To quantify the bias and noise in behavioral scientists' predictions across diverse behavioral experiments.
- To evaluate whether simple statistical models, including linear interpolations and empirical Bayes estimators, outperform expert forecasts.
- To examine the replicability of predictions in behavioral science and the credibility of expert judgment in forecasting outcomes.
Proposed method
- Compared behavioral scientists' predictions against a null model assuming no effect of interventions (i.e., zero predicted change in behavior).
- Applied nonparametric and parametric empirical Bayes estimators to estimate posterior distributions of true treatment effects, accounting for uncertainty.
- Used Bayesian bootstrap with Dirichlet priors to cluster standard errors at the treatment and behavioral scientist levels, improving robustness of comparisons.
- Employed Gibbs sampling in linear regression models to estimate replication probabilities, incorporating uncertainty in effect size estimates.
- Used comparative risk analysis to compute the probability that expert predictions are less accurate than model predictions, using bootstrapped sampling.
- Applied linear models with OLS and variance-inflated error structures to predict replication likelihood based on original study effect sizes.
Experimental results
Research questions
- RQ1Do behavioral scientists predict behavioral outcomes more accurately than simple heuristics such as 'interventions have no effect'?
- RQ2How does the accuracy of behavioral scientists' predictions compare to random chance across multiple behavioral experiments?
- RQ3To what extent are behavioral scientists' predictions biased, and do they systematically overestimate the effectiveness of interventions?
- RQ4Can simple statistical models, such as linear interpolation or empirical Bayes estimators, outperform expert forecasts in behavioral science?
- RQ5What is the probability that expert predictions are less accurate than model-based predictions, given uncertainty in treatment effects?
Key findings
- Behavioral scientists' predictions were no more accurate than random chance and often worse, particularly in predicting intervention effects.
- The null model predicting zero effect for all interventions outperformed behavioral scientists in all five studies analyzed.
- Experts systematically overestimated the effectiveness of behavioral interventions, showing a consistent positive bias in their forecasts.
- Prediction accuracy was low and noisy, with high variability across individual scientists, even among those from top institutions.
- Simple models such as linear interpolation between known data points performed significantly better than expert predictions in the effort study.
- The probability that expert predictions were less accurate than model predictions exceeded 80% in multiple studies, indicating strong evidence for model superiority.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.