Skip to main content
QUICK REVIEW

[论文解读] A No-Free-Lunch Theorem for MultiTask Learning

Steve Hanneke, Samory Kpotufe|arXiv (Cornell University)|Jun 29, 2020
Domain Adaptation and Few-Shot Learning被引用 4
一句话总结

本文為多任務學習建立了一個無免費午餐定理,證明當所有任務共享同一組最佳分類器時,除非擁有額外的分佈資訊(例如,根據任務與目標的接近程度對其排序),否則任何自適應演算法都無法保證隨著任務數 $N$ 增加而提升泛化率。即使最佳分類器完全相同,最佳聚合方式也可能排除部分來源資料,且在缺乏此類資訊的情況下,收斂率在根本上受雜訊水準 $\beta$ 所限制。

ABSTRACT

Multitask learning and related areas such as multi-source domain adaptation address modern settings where datasets from $N$ related distributions $\{P_t\}$ are to be combined towards improving performance on any single such distribution ${\cal D}$. A perplexing fact remains in the evolving theory on the subject: while we would hope for performance bounds that account for the contribution from multiple tasks, the vast majority of analyses result in bounds that improve at best in the number $n$ of samples per task, but most often do not improve in $N$. As such, it might seem at first that the distributional settings or aggregation procedures considered in such analyses might be somehow unfavorable; however, as we show, the picture happens to be more nuanced, with interestingly hard regimes that might appear otherwise favorable. In particular, we consider a seemingly favorable classification scenario where all tasks $P_t$ share a common optimal classifier $h^*,$ and which can be shown to admit a broad range of regimes with improved oracle rates in terms of $N$ and $n$. Some of our main results are as follows: $\bullet$ We show that, even though such regimes admit minimax rates accounting for both $n$ and $N$, no adaptive algorithm exists; that is, without access to distributional information, no algorithm can guarantee rates that improve with large $N$ for $n$ fixed. $\bullet$ With a bit of additional information, namely, a ranking of tasks $\{P_t\}$ according to their distance to a target ${\cal D}$, a simple rank-based procedure can achieve near optimal aggregations of tasks' datasets, despite a search space exponential in $N$. Interestingly, the optimal aggregation might exclude certain tasks, even though they all share the same $h^*$.

研究动机与目标

  • 探討所有任務共享同一最佳分類器時,多任務學習的根本限制。
  • 確定自適應演算法是否能在任務數 $N$ 和每項任務的樣本數 $n$ 增加時,實現更佳的泛化率。
  • 識別在何種條件下,多個資料集的聚合可帶來優於單一資料集學習的表現。
  • 探討分佈資訊(例如,根據與目標分佈的距離對來源進行排序)在實現自適應改進中的角色。

提出的方法

  • 建立一個包含 $N$ 個相關分佈 $P_t$ 的分類設定,所有分佈共享同一最佳分類器 $h^*$,以分析最小最大率。
  • 推導過剩風險的最小最大上下界,顯示在適當條件下,若擁有 $N$ 和 $n$,則最佳者率會提升。
  • 引入雜訊參數 $\beta \in [0,1)$ 以量化標籤模糊性(例如,Massart 雜訊在 $\beta=1$ 時),此參數決定學習的根本限制。
  • 證明,若無分佈結構的先驗知識,則無任何自適應程序能實現優於 $n^{-1/(2-\beta)}$ 的率。
  • 提出一種基於排序的聚合程序,利用來源與目標分佈距離的排序,即使在 $2^N$ 種可能聚合中搜尋空間呈指數級增長,仍能達成近乎最佳率。
  • 運用反濃度論證與充分統計量(例如,正負標籤的計數)來建立在對抗性構造下過剩風險的下界。

实验结果

研究问题

  • RQ1當所有任務共享同一最佳分類器時,自適應多任務學習演算法是否能實現隨 $N$ 和 $n$ 增加而改善的泛化率?
  • RQ2當缺乏分佈資訊時,性能提升是否存在根本限制?
  • RQ3在存在指數級可能聚合方式的情況下,何種條件可使簡單的基於排序的聚合程序達成近乎最佳表現?
  • RQ4為何多任務學習中的典型理論界通常無法反映 $N$ 增加時的改進,即使在共享最佳分類器的有利設定下?
  • RQ5雜訊水準 $\beta$ 在決定自適應多任務學習根本限制中扮演何種角色?

主要发现

  • 即使所有任務共享同一最佳分類器 $h^*$,任何自適應演算法也無法保證其率優於 $n^{-1/(2-\beta)}$(當 $\beta < 1$ 時),顯示存在根本性的無免費午餐限制。
  • 在 $\beta = 1$(即 Massart 雜訊)時,簡單合併所有資料集可達近乎最小最大最佳率,且在共享 $h^*$ 的假設下,其表現隨 $N$ 和 $n$ 而提升。
  • 即使所有 $h^*$ 完全相同,最佳聚合方式仍可能排除某些來源資料集,顯示並非所有資料都同等貢獻。
  • 根據來源與目標分佈距離的排序,可使簡單的基於排序程序達成近乎最佳的聚合率,即使搜尋空間大小為 $2^N$(呈指數級增長)。
  • 本文建立了一個下界:任何分類器未能達成低於 $c_0 \cdot n^{-1/(2-\beta)}$ 風險的機率至少為 $\frac{1}{12} \cdot \frac{1}{96} \cdot \frac{1}{84} \approx 1.06 \times 10^{-4}$,從而證明更佳自適應率的不可能性。
  • 即使任務數 $N$ 較大且每項任務的樣本數 $n$ 固定,此結果仍成立,顯示若無先驗結構知識,$N$ 的改進無法達成。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。