Skip to main content
QUICK REVIEW

[論文レビュー] A Pseudo-Metric between Probability Distributions based on Depth-Trimmed Regions

Guillaume Staerman, Pavlo Mozharovskyi|arXiv (Cornell University)|Mar 23, 2021
Advanced Statistical Methods and Models被引用数 5
ひとこと要約

本稿では、データデプスを用いて多次元順序統計量(デプス・トリムド領域)を定義することで、$\mathbb{R}^d$ 内の連続的確率分布を比較するための新規な擬似距離を提案する。この距離は、2つの分布に由来するデプスに基づく領域間の平均ハウスドルフ距離を計算し、ロバスト性、アフィン変換不変性、線形時間近似の特徴を有し、テキスト要約評価タスクにおいて、MMD や Wasserstein、BertScore より優れた性能を示す。

ABSTRACT

The design of a metric between probability distributions is a longstanding problem motivated by numerous applications in Machine Learning. Focusing on continuous probability distributions on the Euclidean space $\\mathbb{R}^d$, we introduce a novel pseudo-metric between probability distributions by leveraging the extension of univariate quantiles to multivariate spaces. Data depth is a nonparametric statistical tool that measures the centrality of any element $x\\in\\mathbb{R}^d$ with respect to (w.r.t.) a probability distribution or a data set. It is a natural median-oriented extension of the cumulative distribution function (cdf) to the multivariate case. Thus, its upper-level sets -- the depth-trimmed regions -- give rise to a definition of multivariate quantiles. The new pseudo-metric relies on the average of the Hausdorff distance between the depth-based quantile regions w.r.t. each distribution. Its good behavior w.r.t. major transformation groups, as well as its ability to factor out translations, are depicted. Robustness, an appealing feature of this pseudo-metric, is studied through the finite sample breakdown point. Moreover, we propose an efficient approximation method with linear time complexity w.r.t. the size of the data set and its dimension. The quality of this approximation as well as the performance of the proposed approach are illustrated in numerical experiments.

研究の動機と目的

  • 非重複サポートを持つ確率分布間の、ロバストで幾何学的に意味のある距離を定義する課題に対処すること。
  • データデプスを多次元的な分位数の拡張として活用し、分布比較のためのデプス・トリムド領域を構築すること。
  • アフィン変換に対して不変であり、特に高次元設定においても汚染に対してロバストであることを保証すること。
  • 機械学習応用における実用的導入を可能にする、効率的な線形時間近似手法を開発すること。
  • 自動自然言語生成評価、特にテキスト要約タスクにおける本メトリクスの性能を評価すること。

提案手法

  • データデプスの上位集合(デプス・トリムド領域)を用いて多次元分位数を定義し、分布に対する点の中心性の順序を決定する。
  • 2つの確率分布に由来するデプス・トリムド領域間の平均ハウスドルフ距離として擬似距離を構築する。
  • 各分布に対して、単体的、Oja、ゾノイドデプスなどのデプス関数の族を用いてデプス・トリムド領域を定義する。
  • 有限の分位数レベルグリッド上でデプス値をサンプリングし、距離を計算することで、線形時間近似を提案する。
  • アフィン変換に対して不変であり、位置シフトを要因として除去するように設計され、ロバスト性が向上する。
  • 有限標本ブレイクダウン・ポイントを用いた理論的性質(ロバスト性)の分析を実施する。

実験結果

リサーチクエスチョン

  • RQ1データデプスを用いて、分布比較に意味を持つ多次元的分位数の拡張を定義できるか?
  • RQ2提案された擬似距離がアフィン変換および平行移動に対して望ましい不変性を示すか?
  • RQ3有限標本ブレイクダウン・ポイントで測定した場合、データ汚染に対してどの程度ロバストか?
  • RQ4理論的性質を保持しつつ、効率的で線形時間の近似を構築できるか?
  • RQ5既存のメトリクス(例:MMD、Wasserstein、BertScore)と比較して、テキスト生成品質の評価において本メトリクスはどの程度優れているか?

主な発見

  • 提案された擬似距離 $DR_{p,\varepsilon}$ は、WebNLG 2020ベンチマークにおいて人間の判断との相関が最も高く、抽出的要約システムでは MoverScore や BertScore を上回った。
  • 要約タスクにおいて、$DR_{p,\varepsilon}$ は抽出的システムにおける人間の判断とのスピアマン相関が 91.5 を達成し、Wasserstein(74.2)や MMD(75.6)を上回った。
  • 有限標本ブレイクダウン・ポイントにより、データ汚染に対して強いロバスト性が確認された。
  • 線形時間近似手法は、計算コストを顕著に削減しながらも高い精度を維持し、大規模データセットへのスケーラビリティを実現した。
  • アフィン変換に対して不変であり、平行移動効果を効果的に要因として除去するため、実世界の応用において安定性が向上した。
  • 実験的結果から、$DR_{p,\varepsilon}$ は抽出的および要約的要約評価において、Sliced-Wasserstein や MMD よりも優れた性能を示し、特に Kendall の $\tau$ およびピアソンの $r$ において顕著な優位性を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。