Skip to main content
QUICK REVIEW

[论文解读] Empirical or Invariant Risk Minimization? A Sample Complexity Perspective

Kartik Ahuja, Jun Wang|arXiv (Cornell University)|Oct 30, 2020
Statistical Methods and Inference参考文献 29被引用 15
一句话总结

本文從樣本複雜度的視角比較了經驗風險最小化(ERM)與不變風險最小化(IRM),顯示在混雜變數或反因果變數偏移的情境下,IRM 在分布外泛化表現優於 ERM —— 解決方案與理想預測器的差距在 O(√ε) 內,樣本複雜度為 O(1/ε²),且與模型複雜度呈多項式關係。相比之下,ERM 在此類設定下仍存在漸近偏差,而在協變量偏移下兩者表現相近。

ABSTRACT

Recently, invariant risk minimization (IRM) was proposed as a promising solution to address out-of-distribution (OOD) generalization. However, it is unclear when IRM should be preferred over the widely-employed empirical risk minimization (ERM) framework. In this work, we analyze both these frameworks from the perspective of sample complexity, thus taking a firm step towards answering this important question. We find that depending on the type of data generation mechanism, the two approaches might have very different finite sample and asymptotic behavior. For example, in the covariate shift setting we see that the two approaches not only arrive at the same asymptotic solution, but also have similar finite sample behavior with no clear winner. For other distribution shifts such as those involving confounders or anti-causal variables, however, the two approaches arrive at different asymptotic solutions where IRM is guaranteed to be close to the desired OOD solutions in the finite sample regime, while ERM is biased even asymptotically. We further investigate how different factors -- the number of environments, complexity of the model, and IRM penalty weight -- impact the sample complexity of IRM in relation to its distance from the OOD solutions

研究动机与目标

  • 解決「在何種情況下應優先選擇 IRM 而非 ERM 以實現分布外泛化」這一開放問題。
  • 分析不同類型分佈偏移下 ERM 與 IRM 的有限樣本與漸近行為。
  • 量化 IRM 相對於真實不變預測器在多項式生成模型中的樣本複雜度。
  • 識別即使在無限數據下 ERM 仍存在漸近偏差的條件。
  • 比較 ERM 與 IRM 在多種分佈偏移機制(包括協變量偏移、混雜變數與反因果變數)下的表現。

提出的方法

  • 透過不變性條件形式化分佈偏移:P(Y|Φ*(X)) 在各環境中保持不變。
  • 在具有線性與多項式生成機制的結構方程模型下分析 ERM 與 IRM。
  • 推導 IRM 的樣本複雜度邊界,顯示其與 IRM 約束中鬆弛參數 ε 的關係為 O(1/ε²)。
  • 利用梯度分析證明當矩陣 ρ̄ 為滿秩時,ERM 存在漸近偏差。
  • 在 IRM 中採用極小-極大優化以強制跨環境的不變性,最小化最壞情況風險。
  • 應用來自優化與高維統計的理論工具,以界定與理想不變預測器的偏離程度。

实验结果

研究问题

  • RQ1在何種類型的分佈偏移下,IRM 在樣本複雜度與有限樣本表現上優於 ERM?
  • RQ2當資料生成過程中存在混雜變數或反因果變數時,ERM 是否為漸近無偏?
  • RQ3IRM 的樣本複雜度如何隨其約束中的鬆弛參數 ε 及模型類別複雜度而變化?
  • RQ4何種條件會導致 ERM 即使在無限數據下仍存在偏差,而 IRM 又如何避免此問題?
  • RQ5環境數量、模型複雜度與 IRM 損失權重如何共同影響 IRM 對不變預測器的收斂性?

主要发现

  • 在協變量偏移(Φ* = identity)下,ERM 與 IRM 會達到相同的漸近解,且有限樣本表現相近,無明顯優劣之分。
  • 在混雜變數或反因果變數偏移情境下(Φ* ≠ identity),ERM 存在漸近偏差,即使擁有無限數據也無法收斂至真實的不變預測器。
  • IRM 可達成與理想不變預測器差距在 O(√ε) 以內的解,樣本複雜度對多項式模型而言為 O(1/ε²)。
  • IRM 的樣本複雜度隨模型類別複雜度呈多項式增長,而 ERM 在非協變量偏移設定下,其偏差不因資料量增加而消除。
  • ERM 的偏差來自矩陣 ρ̄ = [𝔼[εᵉXᵉ]] 的秩至少為一,此條件在非退化分佈下通常成立。
  • IRM 的極小-極大優化在弱正則性條件下可確保收斂至不變預測器,而 ERM 則無法保證此不變性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。