Skip to main content
QUICK REVIEW

[Paper Review] Evaluating General-Purpose AI with Psychometrics

Xiting Wang, Liming Jiang|arXiv (Cornell University)|Oct 25, 2023
Explainable Artificial Intelligence (XAI)4 citations
TL;DR

This paper proposes a shift from task-oriented to construct-oriented evaluation of general-purpose AI using psychometric principles to identify and measure latent cognitive constructs underlying AI performance. By applying psychometrics—originally developed for human intelligence—it enables predictive, explanatory, and reliable evaluation of AI across unforeseen tasks, offering a scientific foundation for assessing versatility and capability beyond predefined benchmarks.

ABSTRACT

Comprehensive and accurate evaluation of general-purpose AI systems such as large language models allows for effective mitigation of their risks and deepened understanding of their capabilities. Current evaluation methodology, mostly based on benchmarks of specific tasks, falls short of adequately assessing these versatile AI systems, as present techniques lack a scientific foundation for predicting their performance on unforeseen tasks and explaining their varying performance on specific task items or user inputs. Moreover, existing benchmarks of specific tasks raise growing concerns about their reliability and validity. To tackle these challenges, we suggest transitioning from task-oriented evaluation to construct-oriented evaluation. Psychometrics, the science of psychological measurement, provides a rigorous methodology for identifying and measuring the latent constructs that underlie performance across multiple tasks. We discuss its merits, warn against potential pitfalls, and propose a framework to put it into practice. Finally, we explore future opportunities of integrating psychometrics with the evaluation of general-purpose AI systems.

Motivation & Objective

  • To address the limitations of current task-oriented benchmarks in evaluating general-purpose AI systems.
  • To propose a construct-oriented evaluation framework grounded in psychometric science for measuring latent AI capabilities.
  • To improve predictive power, explanatory power, and quality assurance in AI evaluation.
  • To guide AI selection, training, and integration based on measurable latent constructs.
  • To identify and mitigate risks associated with unanticipated AI behaviors in real-world applications.

Proposed method

  • Adopting psychometric principles to define and measure latent constructs underlying AI behavior, analogous to cognitive abilities in humans.
  • Using empirical data from diverse AI outputs to infer and validate constructs through statistical modeling.
  • Designing a three-stage evaluation framework: selection, training, and validation, informed by construct measurement.
  • Extending psychometric techniques such as item response theory and confirmatory factor analysis to AI evaluation.
  • Integrating feedback from construct measurement to optimize AI training and improve performance on targeted constructs.
  • Reinterpreting human-centric psychometric constructs for non-human AI systems while accounting for prompt sensitivity and model variability.
Figure 1: Comparison of the task-oriented paradigm for AI evaluation and psychometrics.
Figure 1: Comparison of the task-oriented paradigm for AI evaluation and psychometrics.

Experimental results

Research questions

  • RQ1How can psychometric methods improve the predictive and explanatory power of AI evaluation beyond task-specific benchmarks?
  • RQ2What latent constructs underlie the versatile performance of general-purpose AI systems, and how can they be reliably measured?
  • RQ3How can psychometric evaluation support the selection and training of AI systems for high-stakes applications?
  • RQ4What are the risks and limitations of directly applying human psychometric tests to evaluate AI systems?
  • RQ5How can psychometric evaluation ensure the reliability and validity of AI assessments in real-world integration?

Key findings

  • Psychometric evaluation offers superior predictive power by identifying latent constructs that generalize across unseen tasks.
  • Construct-oriented evaluation enables deeper explanation of AI performance variations across different inputs or prompts.
  • The framework supports quality-assured testing by ensuring reliability and validity in AI evaluation, reducing bias and inconsistency.
  • Psychometrics can guide the selection of high-potential AI systems for specific roles, such as legal assistants, before extensive training.
  • The method reveals fundamental limitations in AI systems, such as lack of critical thinking, through systematic construct measurement.
  • The integration of psychometrics into AI evaluation paves the way for more responsible, transparent, and scientifically grounded AI development practices.
Figure 2: A framework for construct-oriented evaluation grounded in psychometrics, illustrated by an example of evaluating a general medical AI assistant, with exemplary psychometric techniques at each stage.
Figure 2: A framework for construct-oriented evaluation grounded in psychometrics, illustrated by an example of evaluating a general medical AI assistant, with exemplary psychometric techniques at each stage.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.