[Paper Review] Your 2 is My 1, Your 3 is My 9: Handling Arbitrary Miscalibrations in Ratings
This paper proposes a novel class of estimators that leverage cardinal ratings with arbitrary, unknown miscalibrations—without assuming linear or parametric biases—yet strictly and uniformly outperform any estimator relying only on induced rankings. Drawing on empirical Bayes and Stein's shrinkage, the method exploits cardinal score structure to achieve superior performance in A/B testing and ranking, challenging the long-held belief that only ordinal rankings are useful under miscalibration.
Cardinal scores (numeric ratings) collected from people are well known to suffer from miscalibrations. A popular approach to address this issue is to assume simplistic models of miscalibration (such as linear biases) to de-bias the scores. This approach, however, often fares poorly because people's miscalibrations are typically far more complex and not well understood. In the absence of simplifying assumptions on the miscalibration, it is widely believed by the crowdsourcing community that the only useful information in the cardinal scores is the induced ranking. In this paper, inspired by the framework of Stein's shrinkage, empirical Bayes, and the classic two-envelope problem, we contest this widespread belief. Specifically, we consider cardinal scores with arbitrary (or even adversarially chosen) miscalibrations which are only required to be consistent with the induced ranking. We design estimators which despite making no assumptions on the miscalibration, strictly and uniformly outperform all possible estimators that rely on only the ranking. Our estimators are flexible in that they can be used as a plug-in for a variety of applications, and we provide a proof-of-concept for A/B testing and ranking. Our results thus provide novel insights in the eternal debate between cardinal and ordinal data.
Motivation & Objective
- To challenge the widely held belief that cardinal ratings are only useful in terms of induced rankings when miscalibrations are unknown or arbitrary.
- To design estimators that use cardinal scores but make no assumptions about the form of miscalibration, even if adversarially chosen.
- To demonstrate that cardinal scores contain strictly more information than rankings alone under arbitrary miscalibrations.
- To provide a plug-in framework for improving existing ranking-based algorithms in applications like A/B testing and item ranking.
- To establish theoretical guarantees showing uniform superiority of cardinal-based estimators over any ranking-only alternative.
Proposed method
- Formalizes miscalibration as an unknown monotonic transformation from true values to reported scores, with no parametric assumptions.
- Applies empirical Bayes and Stein's shrinkage framework to construct estimators that shrink or adjust scores toward a common reference, improving estimation under uncertainty.
- Designs a canonical estimator based on a weighted average of scores, where the weights are derived from the observed ranking structure and empirical distributions.
- Uses the two-envelope problem analogy to justify the design of randomized estimators that outperform deterministic ranking-based strategies.
- Derives theoretical bounds showing that the proposed cardinal estimators uniformly dominate all ranking-based estimators in terms of success probability.
- Proves that optimal performance is achieved only when estimators respect the topological ordering induced by the observed rankings, and shows that cardinal estimators can satisfy this condition with higher probability.
Experimental results
Research questions
- RQ1Can cardinal ratings with arbitrary miscalibrations yield better estimation than ranking-based methods, even without modeling the miscalibration?
- RQ2Is it possible to construct estimators that outperform all ranking-only estimators uniformly, without assuming any parametric form of miscalibration?
- RQ3Can the structure of cardinal scores be exploited to improve A/B testing and ranking tasks when reviewer biases are unknown and potentially adversarial?
- RQ4What is the theoretical relationship between cardinal scores and rankings under arbitrary monotonic miscalibrations?
- RQ5How can empirical Bayes and shrinkage techniques be adapted to handle miscalibrated ratings in a way that preserves or improves estimation accuracy?
Key findings
- The proposed cardinal estimators strictly and uniformly outperform any estimator that relies solely on the induced ranking, even when miscalibrations are adversarially chosen.
- The success probability of the proposed estimators exceeds that of any ranking-only estimator for every possible configuration of miscalibrations and true rankings.
- The optimal ranking-based estimator must always output a topological ordering consistent with the observed ranking; the proposed cardinal estimators achieve this with higher probability.
- The theoretical analysis proves that the probability of correct estimation using cardinal scores is bounded below by a value strictly greater than the maximum achievable by any ranking-only method.
- The method is robust to discrete rating scales, with only minor modifications needed to handle ties, and extends naturally to practical settings like peer review.
- The framework enables plug-in improvements for existing A/B testing and ranking algorithms, offering a practical path to enhanced performance without re-engineering the full pipeline.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.