[Paper Review] A Model of Artificial Jagged Intelligence
This paper presents an economic model of AI jaggedness (AJI), where local reliability varies across tasks and adoption depends on discoverability, calibration, and mastery, not just average performance.
Generative AI systems often display highly uneven performance across tasks that appear ``nearby'': they can be excellent on one prompt and confidently wrong on another with only small changes in wording or context. We call this phenomenon Artificial Jagged Intelligence (AJI). This paper develops a tractable economic model of AJI that treats adoption as an information problem: users care about \emph{local} reliability, but typically observe only coarse, global quality signals. In a baseline one-dimensional landscape, truth is a rough Brownian process, and the model ``knows'' scattered points drawn from a Poisson process. The model interpolates optimally, and the local error is measured by posterior variance. We derive an adoption threshold for a blind user, show that experienced errors are amplified by the inspection paradox, and interpret scaling laws as denser coverage that improves average quality without eliminating jaggedness. We then study mastery and calibration: a calibrated user who can condition on local uncertainty enjoys positive expected value even in domains that fail the blind adoption test. Modelling mastery as learning a reliability map via Gaussian process regression yields a learning-rate bound driven by information gain, clarifying when discovering ``where the model works'' is slow. Finally, we study how scaling interacts with discoverability: when calibrated signals and user mastery accelerate the harvesting of scale improvements, and when opacity can make gains from scaling effectively invisible.
Motivation & Objective
- Explain how local reliability heterogeneity affects AI adoption and productivity in knowledge-work settings.
- Introduce a tractable baseline model combining a Poisson knowledge-point process with a Brownian truth landscape.
- Derive adoption thresholds under blind use and analyze the role of inspection paradox in experienced errors.
- Study how scaling (denser knowledge coverage) and calibration alter welfare and degenerate effects of jaggedness.
- Expose how mastery and interface design complement raw model improvements for productive human–AI collaboration.
Proposed method
- Model the knowledge points as a Poisson point process with intensity lambda to represent coverage density.
- Represent the truth landscape Y(x) as a Brownian motion to generate rough interpolation risk between knowledge points.
- Compute the posterior variance sigma^2(x) for interpolation between adjacent knowledge points. sigma^2(x) = (x-x_i)(x_{i+1}-x)/(x_{i+1}-x_i).
- Define user payoff U(x) = 1 - sigma^2(x)/q and analyze adoption under blind use (outside option 0).
- Derive an adoption threshold under blind adoption: q >= 1/(3 lambda) (equivalently R >= 1 with R = 3 lambda q).
- Introduce calibration as a benchmark where users observe sigma^2(x) and show positive value U_C(R) that depends on R.
Experimental results
Research questions
- RQ1How does local, task-level reliability heterogeneity affect AI adoption and welfare under opacity?
- RQ2What is the adoption threshold when users rely on AI without task-specific reliability signals?
- RQ3How does scaling (denser coverage, larger lambda) alter expected error and jaggedness?
- RQ4How does calibration or mastery change the value of AI assistance and its adoption under AJI?
- RQ5What are the complementarities or substitutions between scaling, calibration, and mastery in improving productivity?
Key findings
- Blind adoption can be optimal only when the reliability index R = 3 lambda q is at least 1.
- Scaling (increasing lambda) reduces local posterior variance, but jaggedness persists in shape.
- Calibration converts jaggedness into an option value, yielding positive welfare gains relative to blind use.
- Mastery yields learning rates governed by information gain; learning where the model works can be slow in high-dimensional spaces.
- Complementarities exist: scaling and calibration/mastery can be substitutes or complements depending on the adoption threshold; interface design can aid discoverability without full model improvements.
- The model highlights an inspection paradox: experienced error is amplified because users spend more time in longer gaps between knowledge points.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.