Skip to main content
QUICK REVIEW

[論文レビュー] Hockey Player Performance via Regularized Logistic Regression

Robert B. Gramacy, Matt Taddy|arXiv (Cornell University)|Dec 21, 2016
Sports Analytics and Performance参考文献 19被引用数 3
ひとこと要約

本論文は、チームメートや相手選手の影響を同時にモデル化することで、フィールド上での得点貢献度を個別に推定する正則化付きロジスティック回帰モデルを提案する。高次元のゲームデータに罰則付き尤度推定を適用することで、各選手のチーム全体のノイズを超えた影響を分離し、従来のプラスマイナスよりもより正確で安定したパフォーマンス評価を可能にする。

ABSTRACT

A hockey player's plus-minus measures the difference between goals scored by and against that player's team while the player was on the ice. This measures only a marginal effect, failing to account for the influence of the others he is playing with and against. A better approach would be to jointly model the effects of all players, and any other confounding information, in order to infer a partial effect for this individual: his influence on the box score regardless of who else is on the ice. This chapter describes and illustrates a simple algorithm for recovering such partial effects. There are two main ingredients. First, we provide a logistic regression model that can predict which team has scored a given goal as a function of who was on the ice, what teams were playing, and details of the game situation (e.g. full-strength or power-play). Since the resulting model is so high dimensional that standard maximum likelihood estimation techniques fail, our second ingredient is a scheme for regularized estimation. This adds a penalty to the objective that favors parsimonious models and stabilizes estimation. Such techniques have proven useful in fields from genetics to finance over the past two decades, and have demonstrated an impressive ability to gracefully handle large and highly imbalanced data sets. The latest software packages accompanying this new methodology -- which exploit parallel computing environments, sparse matrices, and other features of modern data structures -- are widely available and make it straightforward for interested analysts to explore their own models of player contribution.

研究の動機と目的

  • ホッケーにおける個別選手の貢献度を分離する際の従来のプラスマイナスの限界を解消すること。
  • チームメート、相手選手、およびゲームの文脈を考慮して、すべての選手が得点結果に与える連合的影響をモデル化すること。
  • 高次元で不均衡なホッケーのデータに対して、安定的で単純な推定手法を開発すること。
  • 現代の正則化技術を用いて、プラスマイナスのような単純な指標の代わりに統計的に厳密な代替手法を提供すること。
  • アクセスしやすいソフトウェアと計算効率を活用して、実用的な選手評価モデルの応用を可能にすること。

提案手法

  • 得点を記録したチームがどのチームであったかを結果とするロジスティック回帰モデルを構築し、選手のフィールド上での出場状況、チーム識別子、ゲームの状況(例:ペナルティキック、エブンスティール)を入力とする。
  • 標準的最尤推定が失敗する高次元パrameter空間において、推定を安定化させるために正則化(例:L1またはL2ペナルティ)を用いる。
  • スパース行列表現と並列計算を活用して、大規模データを効率的に処理する。
  • 罰則付き対数尤度の目的関数を最適化し、未知のデータに一般化しやすい単純なモデルを優遇する。
  • 各選手の部分的効果を推定し、チームの文脈とは独立して得点貢献度を表す。
  • 正則化と高性能コンピューティングをサポートする現代の統計ソフトウェアパッケージを活用して、モデルの適合を実現する。

実験結果

リサーチクエスチョン

  • RQ1ホッケーにおける個別選手の貢献度は、チームや相手選手の影響からどのように分離可能か?
  • RQ2ロジスティック回帰を用いた連合的モデリングアプローチは、プラスマイナスのような単純指標よりも選手の影響をよりよく推定できるか?
  • RQ3正則化は、高次元かつスパースなホッケーのデータにおいて推定の安定性をどのように向上させるか?
  • RQ4ゲームの文脈(例:ペナルティキック、ショートハンデッド)を、選手パフォーマンスモデルに意味的に組み込める範囲はどの程度か?
  • RQ5スケーラブルな計算手法は、高価なリソースを持たないアナリストにとっても、高度な選手評価モデルを実用可能にするか?

主な発見

  • 提案されたモデルは、チームメートや相手選手の影響を考慮することで、従来のプラスマイナスよりも正確なフィールド上での影響測定を可能にする。
  • 正則化により、過学習や同定不能性のため標準的最尤推定が失敗する高次元設定でも推定が安定する。
  • この手法は、ショートハンデッドゴールのようなレアイベントを含む不均衡なデータを、変換なしに効果的に処理できる。
  • スパース行列と並列計算の活用により、大規模データセットでも効率的なモデル適合が可能となり、実世界の分析に実用的である。
  • 得られた部分的効果推定値は解釈可能で頑健であり、ヒューリスティックなパフォーマンス指標の代替としての原則的根拠を持つ。
  • 現代のソフトウェアツールのおかげで、研究者やアナリストが独立してこのようなモデルを実装・拡張することが現実可能になった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。