Skip to main content
QUICK REVIEW

[論文レビュー] Using Machine Learning to Test Causal Hypotheses in Conjoint Analysis

Dae-Woong Ham, Kosuke Imai|arXiv (Cornell University)|Jan 20, 2022
Economic and Environmental Valuation参考文献 34被引用数 7
ひとこと要約

本稿は、交差分析における因果推論のための条件付きランダム化検定(CRT)ベースの手法を提案する。この手法は仮定フリーであり、機械学習を用いて複雑な交互作用を検出可能である。ランダム化に基づく推論を採用することで、モデルの誤指定を回避しつつ、高次元設定下でも強力で正確なp値を提供する。従来の平均的マージナル成分効果(AMCE)手法に比べ、顕著な交互作用の検出において優れている。

ABSTRACT

Conjoint analysis is a popular experimental design used to measure multidimensional preferences. Researchers examine how varying a factor of interest, while controlling for other relevant factors, influences decision-making. Currently, there exist two methodological approaches to analyzing data from a conjoint experiment. The first focuses on estimating the average marginal effects of each factor while averaging over the other factors. Although this allows for straightforward design-based estimation, the results critically depend on the distribution of other factors and how interaction effects are aggregated. An alternative model-based approach can compute various quantities of interest, but requires researchers to correctly specify the model, a challenging task for conjoint analysis with many factors and possible interactions. In addition, a commonly used logistic regression has poor statistical properties even with a moderate number of factors when incorporating interactions. We propose a new hypothesis testing approach based on the conditional randomization test to answer the most fundamental question of conjoint analysis: Does a factor of interest matter in any way given the other factors? Our methodology is solely based on the randomization of factors, and hence is free from assumptions. Yet, it allows researchers to use any test statistic, including those based on complex machine learning algorithms. As a result, we are able to combine the strengths of the existing design-based and model-based approaches. We illustrate the proposed methodology through conjoint analysis of immigration preferences and political candidate evaluation. We also extend the proposed approach to test for regularity assumptions commonly used in conjoint analysis. An open-source software package is available for implementing the proposed methodology.

研究の動機と目的

  • 従来のAMCEベースの手法が、他の要因の平均化によって顕著な交互作用を隠してしまうという限界を是正すること。
  • 交差実験における因果効果の仮定フリーな仮説検定を可能にする手法の開発。パラメトリックモデルの仮定に依存しない。
  • 機械学習アルゴリズムを交差データの因果推論に統合し、モデルの誤指定なしに複雑で高次の交互作用を検出可能にする。
  • フレームワークを拡張し、交差実験における正規性仮定(例:プロファイル順序効果なし、疲労効果、持ち越し効果)の検証を可能にする。
  • 研究者が正確なp値を用いて提案手法を実装できる実用的でオープンソースのRパッケージ(CRTConjoint)の提供。

提案手法

  • 本手法は、交差実験のランダム化メカニズムにのみ依存する正確で仮定フリーの仮説検定を実現するため、条件付きランダム化検定(CRT)を採用する。
  • 研究者は、Lassoロジスティック回帰などの複雑な機械学習モデルから得られるテスト統計量を含め、いかなるテスト統計量も使用可能であり、モデル仕様の必要がない。
  • CRTは、設計に条件付けながら、関心のある処置変数の順列を入れ替えることで、帰無仮説の下でのノンシャープ分布を構築する。これにより、鋭い帰無仮説の下で有効な推論が可能になる。
  • 交互作用の検出には、主効果と交互作用項(例:性別と政党所属の間の交互作用)を含むテスト統計量を用い、帰無仮説の下での順列によりp値を計算する。
  • 正規性仮定の検証には、プロファイル順序効果や疲労の違反を反映する適切なテスト統計量を用いることで、アプローチを拡張する。
  • 高次元データと正確なp値をサポートするオープンソースのRパッケージ、CRTConjointを提供する。

実験結果

リサーチクエスチョン

  • RQ1他の要因との交互作用が存在する場合、その関心のある要因がマージナル効果にかかわらず、何らかの形で重要であるかどうか。
  • RQ2機械学習ベースのテスト統計量は、従来のAMCE手法に比べて、交差データにおける顕著な交互作用をより効果的に検出できるか。
  • RQ3交差分析における一般的な正規性仮定(例:プロファイル順序効果なし、疲労)が、現実のデータで破られているか。
  • RQ4交互作用が存在する状況において、CRTベースのアプローチはAMCEベースの推論に比べ、統計的パワーと妥当性の点で優れているか。
  • RQ5CRTは、標準の回帰モデルが見逃す可能性のある高次元の交互作用(例:三重の交互作用)を検出できるか。

主な発見

  • 大統領候補データにおいて、CRTは性別と政党所属の間で顕著な交互作用を検出。HierNetテスト統計量を用いたp値は0.029であり、性別効果が政党に依存することを示唆している。
  • 大統領データにおける最も強い交互作用は、性別と政党所属の間であり、低p値を示しており、性別効果が候補者の政党に依存していることを示唆している。
  • 議会候補者データでは、同じ交互作用のCRTp値は0.029であり、データセットをまたいで一貫した交互作用効果の証拠が得られた。
  • 本手法は、2つの顕著な三重交互作用を同定した。1つは性別、政党所属、回答者の政治的関心を含み、もう1つは性別、政党所属、回答者の自らの政党所属を含む。
  • 移民政策データでは、性別と政党所属の交互作用のCRTp値は0.15であり、中程度の交互作用の証拠を示している。一方、AMCEベースのp値は0.89であり、この効果を検出できていないことを示している。
  • AMCEモデルに主効果と交互作用を追加してもp値は0.40のままだったため、CRTは複雑な交互作用が存在する状況で非ゼロ効果のより強い証拠を提供していることがわかった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。