[Paper Review] Reliability and Comparability of Peer Review Results
This study investigates the reliability and comparability of peer review in academic evaluation by analyzing ratings of research teams at two Belgian universities. It finds that peer review outcomes vary significantly due to inconsistent rating practices, and recommends standardized reference levels and reliability checks to improve consistency, with citation analysis showing mixed alignment with peer ratings depending on subject area and indicator type.
In this paper peer review reliability is investigated based on peer ratings of research teams at two Belgian universities. It is found that outcomes can be substantially influenced by the different ways in which experts attribute ratings. To increase reliability of peer ratings, procedures creating a uniform reference level should be envisaged. One should at least check for signs of low reliability, which can be obtained from an analysis of the outcomes of the peer evaluation itself. The peer review results are compared to outcomes from a citation analysis of publications by the same teams, in subject fields well covered by citation indexes. It is illustrated how, besides reliability, comparability of results depends on the nature of the indicators, on the subject area and on the intrinsic characteristics of the methods. The results further confirm what is currently considered as good practice: the presentation of results for not one but for a series of indicators.
Motivation & Objective
- To assess the reliability of peer review outcomes in evaluating research teams at Belgian universities.
- To investigate how variations in expert rating practices affect the consistency and validity of peer review results.
- To compare peer review outcomes with citation analysis results to evaluate comparability across different evaluation indicators.
- To identify methodological factors influencing reliability and comparability in research evaluation.
- To recommend procedural improvements—such as standardized reference levels and reliability checks—for more consistent peer review practices.
Proposed method
- Collected peer review ratings from experts evaluating research teams at two Belgian universities.
- Analyzed variations in rating patterns across reviewers to assess reliability and consistency.
- Conducted citation analysis on publications by the same research teams using citation indexes in well-covered subject areas.
- Compared peer review results with citation-based indicators to assess alignment and comparability.
- Used statistical analysis of evaluation outcomes to detect signs of low reliability within the review process.
- Evaluated the impact of subject area and indicator type on the comparability of peer review and citation-based results.
Experimental results
Research questions
- RQ1To what extent do peer review outcomes vary due to differences in how experts assign ratings?
- RQ2How comparable are peer review results to citation analysis outcomes for the same research teams?
- RQ3What role do subject area characteristics and indicator types play in the reliability and comparability of evaluation results?
- RQ4Can reliability in peer review be improved through standardized reference levels and internal consistency checks?
- RQ5What indicators best support trustworthy and comparable research evaluation outcomes?
Key findings
- Peer review outcomes showed substantial variability due to inconsistent rating practices among experts, indicating low reliability.
- The study found that different experts applied ratings using divergent reference points, undermining consistency.
- Citation analysis results were only partially aligned with peer review outcomes, with alignment varying by subject area.
- The reliability of peer review was found to be sensitive to the choice of indicators and the intrinsic characteristics of evaluation methods.
- The study confirmed that presenting results across multiple indicators improves the robustness and credibility of evaluations.
- Signs of low reliability in peer review can be detected through internal analysis of evaluation outcomes, supporting the need for such checks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.