[Paper Review] Bayesian Network Models for Adaptive Testing
This paper proposes and evaluates Bayesian network models for computerized adaptive testing (CAT) using data from grammar school mathematics tests. By modeling student knowledge through skill nodes with varying states and incorporating item responses, the models adaptively select questions to efficiently estimate ability; results show that models with three or more skill states outperform simpler ones, while expert models with high complexity underperform due to overfitting and sparsity, suggesting a need for more data or structural simplification.
Computerized adaptive testing (CAT) is an interesting and promising approach to testing human abilities. In our research we use Bayesian networks to create a model of tested humans. We collected data from paper tests performed with grammar school students. In this article we first provide the summary of data used for our experiments. We propose several different Bayesian networks, which we tested and compared by cross-validation. Interesting results were obtained and are discussed in the paper. The analysis has brought a clearer view on the model selection problem. Future research is outlined in the concluding part of the paper.
Motivation & Objective
- To develop Bayesian network models that enable efficient, adaptive assessment of student knowledge in mathematics.
- To evaluate how different model structures—especially skill variable state counts—affect the accuracy and efficiency of adaptive testing.
- To investigate the impact of additional student information (e.g., grades) on model performance during early testing stages.
- To compare the performance of simple, expert, and information-augmented Bayesian network models using real student test data.
- To identify structural and data-related factors that influence model reliability and generalization in adaptive testing.
Proposed method
- Constructed Bayesian networks with hidden skill nodes representing student knowledge in mathematical functions, using data from 281 grammar school students.
- Model structures varied in the number of skill variable states (2, 3, or more) and included Boolean or integer-valued subproblems as evidence nodes.
- Used cross-validation to evaluate model performance, simulating adaptive testing by sequentially selecting questions based on current belief updates.
- Incorporated both binary (correct/incorrect) and numeric (0–4 point) scoring for subproblems to assess sensitivity to response granularity.
- Evaluated models using success ratio and question selection patterns, analyzing conditional probability table (CPT) sparsity and parameter count to detect overfitting.
- Compared models with and without additional student information (e.g., prior grades) to assess its impact on early-stage accuracy.
Experimental results
Research questions
- RQ1How does increasing the number of states in hidden skill variables affect the accuracy of adaptive testing models?
- RQ2How do expert models with complex skill structures compare to simpler models in terms of predictive performance and robustness?
- RQ3To what extent does incorporating prior student information (e.g., grades) improve early-stage estimation accuracy in adaptive testing?
- RQ4What role does model complexity and CPT sparsity play in overfitting and performance degradation, especially with limited data?
- RQ5How do different model structures influence the diversity and consistency of question selection during the testing process?
Key findings
- Models with three or more states for hidden skill variables significantly outperformed those with only two states, particularly during the most critical early stages of testing.
- The expert model, despite its high complexity, performed worse than simpler models, primarily due to overfitting and high CPT sparsity, as indicated by an average sparsity of 0.121 and 81.7 zeros per CPT row.
- Additional student information (e.g., prior grades) improved performance only in the initial testing steps, suggesting limited long-term benefit and potential practical or ethical drawbacks.
- The number of zeros in CPTs (AZT) and sparsity (AS) increased with model complexity, confirming that limited data volume exacerbates overfitting in high-parameter models.
- Question selection patterns showed that expert models exhibited high variability and less consistency across test runs, indicating uncertainty in selection strategies.
- Simpler models with 3-state skill variables achieved better balance between accuracy and stability, suggesting they are more suitable for real-world adaptive testing applications.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.