Skip to main content
QUICK REVIEW

[论文解读] Analysis of Information Transfer from Heterogeneous Sources via Precise High-dimensional Asymptotics

Fan Yang, Hongyang R. Zhang|arXiv (Cornell University)|Oct 22, 2020
Stochastic Gradient Optimization Techniques参考文献 52被引用 4
一句话总结

本文通过共享参数的两层线性神经网络,分析异质任务间的信息迁移,推导出在样本量与特征维度成比例增长时预测风险的精确高维渐近行为。研究识别出正向或负向迁移的精确条件,揭示了在模型偏移下风险曲线的非单调性,并在文本分类任务上验证了渐进式数据添加策略的有效性。

ABSTRACT

We consider the problem of learning -- gaining knowledge from one source task and applying it to a different but related target task. A fundamental question in learning is whether combining the data of both tasks works better than using only the target task's data (equivalently, whether a information transfer happens). We study this question formally in a linear regression setting where a two-layer linear neural network estimator combines both tasks' data. The estimator uses a shared parameter vector for both tasks and exhibits positive or negative information by varying dataset characteristics. We characterize the precise asymptotic limit of the prediction risk of the above estimator when the sample sizes increase with the feature dimension proportionally at fixed ratios. We also show that the asymptotic limit is sufficiently accurate for finite dimensions. Then, we provide the exact condition to determine positive (and negative) information in a random-effect model, leading to several theoretical insights. For example, the risk curve is non-monotone under model shift, thus motivating a learning procedure that progressively adds data from the source task. We validate this procedure's efficiency on text classification tasks with a neural network that applies a shared feature space for both tasks, similar to the above estimator. The main ingredient of the analysis is finding the high-dimensional asymptotic limits of various functions involving the sum of two independent sample covariance matrices with different population covariance matrices, which may be of independent interest.

研究动机与目标

  • 正式确定结合源任务和目标任务数据是否能提升仅使用目标数据时的预测性能。
  • 刻画共享参数两层线性神经网络估计器在高维渐近下的精确预测风险。
  • 在随机效应模型中,识别信息迁移为正或负的精确条件。
  • 提出并验证一种渐进式数据添加学习过程,可适应模型偏移与非单调风险行为。
  • 推导涉及具有不同总体结构的独立样本协方差矩阵之和的函数的高维渐近极限。

提出的方法

  • 在两层线性神经网络估计器中,采用跨源任务与目标任务共享参数向量的线性回归框架,形式化学习问题。
  • 在样本量与特征维度以固定比例同步增长的高维渐近设定下,分析预测风险。
  • 推导涉及两个具有不同总体协方差矩阵的独立样本协方差矩阵之和的函数的精确渐近极限。
  • 利用随机矩阵理论工具,刻画估计器风险在不同数据集特征下的渐近行为。
  • 提出一种渐进式学习过程,根据观测到的非单调风险曲线,逐步添加源任务数据。
  • 在使用共享特征空间的神经网络上,通过文本分类任务验证所提方法,模拟理论估计器的行为。

实验结果

研究问题

  • RQ1在何种条件下,结合源任务与目标任务数据能提升预测性能,优于仅使用目标数据?
  • RQ2当特征数与样本数成比例增长时,共享参数线性估计器的预测风险如何渐近表现?
  • RQ3在具有异质数据的随机效应模型中,什么决定了信息迁移为正或为负?
  • RQ4模型偏移如何影响风险曲线?是否可借此设计更优的学习过程?
  • RQ5所推导的渐近极限在实际有限维设置中,多大程度上准确反映真实行为?

主要发现

  • 在模型偏移下,共享参数估计器的预测风险对源任务数据规模表现出非单调依赖,表明添加数据并不总能提升性能。
  • 根据源任务与目标任务的相对特征结构与样本规模,推导出正向或负向信息迁移的精确条件。
  • 即使在有限维设置中,渐近风险极限也被证明具有高度准确性,验证了理论近似的有效性。
  • 非单调风险曲线启发了一种渐进式数据添加策略,通过避免有害的源数据贡献,提升了泛化性能。
  • 对具有不同总体结构的两个独立样本协方差矩阵之和的理论分析,在随机矩阵理论中具有独立研究价值。
  • 在文本分类任务上的实证验证确认了渐进式学习过程的有效性,其性能优于简单的数据组合方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。