Skip to main content
QUICK REVIEW

[論文レビュー] Does quantification without adjustments work?

Dirk Tasche|arXiv (Cornell University)|Feb 28, 2016
Machine Learning and Data Classification参考文献 13被引用数 4
ひとこと要約

本稿は、分類器がQ測度基準に従って特に定量化の目的で訓練された場合、後処理の調整なしに定量化が機能するかどうかを調査する。Classify & Countは、訓練とターゲットの周辺確率が同一の場合には適切に機能するが、Q測度アプローチは誤較正のリスクを伴い、実用的有用性が訓練とターゲットのクラス周辺確率がほぼ同一である場合に限られることが示された。

ABSTRACT

Classification is the task of predicting the class labels of objects based on the observation of their features. In contrast, quantification has been defined as the task of determining the prevalences of the different sorts of class labels in a target dataset. The simplest approach to quantification is Classify & Count where a classifier is optimised for classification on a training set and applied to the target dataset for the prediction of class labels. In the case of binary quantification, the number of predicted positive labels is then used as an estimate of the prevalence of the positive class in the target dataset. Since the performance of Classify & Count for quantification is known to be inferior its results typically are subject to adjustments. However, some researchers recently have suggested that Classify & Count might actually work without adjustments if it is based on a classifer that was specifically trained for quantification. We discuss the theoretical foundation for this claim and explore its potential and limitations with a numerical example based on the binormal model with equal variances. In order to identify an optimal quantifier in the binormal setting, we introduce the concept of local Bayes optimality. As a side remark, we present a complete proof of a theorem by Ye et al. (2012).

研究の動機と目的

  • 分類器を定量化の目的に特化して訓練した場合、調整なしのClassify & Count定量化の理論的・実証的妥当性を評価すること。
  • 最適な定量器の設計における較正と分類能力の役割を明確にすること。
  • Q測度基準が調整なしの状況で信頼できる定量化性能をもたらすかどうかを調査すること。
  • 特に分散が等しい二正規モデルにおいて、調整なしの定量化が実用的と見なせる条件を同定すること。

提案手法

  • 等分散をもつ二正規モデルを用いて、異なる最適化基準下での定量器性能を解析的に導出し、比較する。
  • 二正規設定における最適な定量器を特定するために、局所ベイズ最適性の概念を導入する。
  • Q測度に異なるβ重みを適用して、較正(NAS*)と真正陽性率(TPR)のバランスをとる。数値実験にはβ=1およびβ=2を用いる。
  • Q測度最適化分類器の性能を、ターゲットデータセットにおける予測された周辺確率と実際の周辺確率の比較によって評価する。
  • 理論的基盤を強化するため、Yeら(2012)の定理の完全な証明を副次的結果として提示する。
  • 固定パラメータμ=0、ν=2、σ=1、P[A]=p=25%を用いて、数値シミュレーションを実施する。

実験結果

リサーチクエスチョン

  • RQ1分類器が定量化の目的に特化して訓練された場合、調整なしのClassify & Count定量化は機能するか?
  • RQ2Q測度基準が調整なしの状況で信頼できる定量化推定をもたらす条件は何か?
  • RQ3Q測度におけるβの選択が、得られる定量器の較正と性能にどのように影響するか?
  • RQ4局所ベイズ最適性で定義される局所最適分類器は、周辺確率推定誤差を最小化する点で他の定量器よりも優れているか?
  • RQ5定量器の性能が、訓練データセットとターゲットデータセットのクラス周辺確率の類似度にどれほど依存するか?

主な発見

  • β=2のQ測度は、訓練とターゲット周辺確率が同一の場合に、完全な性能を達成する局所的最良分類器を特定する。
  • β=1の場合、Q測度最適分類器はu>p=0.25で最大値を示し、較正とTPRのトレードオフが不適切な周辺確率推定をもたらす。
  • Q測度アプローチは誤較正の定量器を生じさせる可能性があり、調整なしのClassify & Countの信頼性を損なう。
  • 調整なしのClassify & Countは、ターゲットデータセットの真正陽性クラス周辺確率が25%の訓練周辺確率とほぼ一致する場合にのみ、良好に機能する。
  • Corollary 2.8により特定されたミニマックス分類器は、β=1の場合、特定の周辺確率領域ではQ測度最適分類器とほぼ同等の性能を示すが、その範囲に限る。
  • 全体として、調整なしの定量化が成功する可能性は、訓練とターゲットデータセットの周辺確率がほぼ同一である場合に限られる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。