Skip to main content
QUICK REVIEW

[论文解读] A General Scoring Rule for Randomized Kernel Approximation with Application to Canonical Correlation Analysis

Yinsong Wang, Shahin Shahrampour|arXiv (Cornell University)|Oct 11, 2019
Gaussian Processes and Bayesian Inference参考文献 32被引用 4
一句话总结

本文提出了一种适用于数据相关随机特征采样的通用评分规则,可提升多种机器学习任务中的核近似性能,尤其聚焦于典型相关分析(CCA)。通过推导出一种基于数据结构以最大化典型相关性的特征采样原则性分布,该方法在 CCA 实验中优于现有技术,同时恢复了已知方法(如杠杆度分数和基于能量的采样)。

ABSTRACT

Random features has been widely used for kernel approximation in large-scale machine learning. A number of recent studies have explored data-dependent sampling of features, modifying the stochastic oracle from which random features are sampled. While proposed techniques in this realm improve the approximation, their application is limited to a specific learning task. In this paper, we propose a general scoring rule for sampling random features, which can be employed for various applications with some adjustments. We first observe that our method can recover a number of data-dependent sampling methods (e.g., leverage scores and energy-based sampling). Then, we restrict our attention to a ubiquitous problem in statistics and machine learning, namely Canonical Correlation Analysis (CCA). We provide a principled guide for finding the distribution maximizing the canonical correlations, resulting in a novel data-dependent method for sampling features. Numerical experiments verify that our algorithm consistently outperforms other sampling techniques in the CCA task.

研究动机与目标

  • 开发一种适用于多种机器学习任务的通用评分规则,用于随机特征采样。
  • 为 CCA 中最大化典型相关性的采样分布提供一种原则性方法。
  • 统一并推广现有的数据相关采样技术,如杠杆度分数和基于能量的采样。
  • 通过优化特征采样分布,提升大规模 CCA 中的核近似精度。
  • 通过 CCA 任务上的数值实验验证该方法的优越性。

提出的方法

  • 作者推导出一种通用评分规则,根据数据结构确定随机特征采样的最优概率。
  • 该方法构建了一种特征分布,以最大化 CCA 中投影子空间之间的期望典型相关性。
  • 通过将现有采样策略嵌入所提出的评分框架中作为特例,实现对现有策略的推广。
  • 采用变分优化方法推导评分规则,以最大化投影子空间之间期望相关性的目标。
  • 通过为对典型相关性贡献更大的特征分配更高采样概率,实现高效采样。
  • 该方法可通过微调评分函数,轻松适配其他基于核的方法。

实验结果

研究问题

  • RQ1如何设计一种通用评分规则,以在多种机器学习任务中改进随机特征采样?
  • RQ2所提出的方法能否恢复已知的数据相关采样技术,如杠杆度分数和基于能量的采样?
  • RQ3在 CCA 中,为最大化典型相关性,最优采样分布是什么?
  • RQ4在 CCA 中,所提出方法与现有采样策略相比性能如何?
  • RQ5该通用评分规则在大规模设置下,能在多大程度上提升核近似精度?

主要发现

  • 所提出的评分规则将已知的数据相关采样方法(如杠杆度分数和基于能量的采样)作为特例恢复。
  • 在 CCA 实验中,该方法获得的典型相关性值高于基线采样技术。
  • 数值结果表明,该方法在多个数据集和设置下均表现出一致的性能提升。
  • 该方法为特征采样提供了一个原则性、统一的框架,其适用范围可扩展至 CCA 以外的其他核方法。
  • 通过聚焦于最能提升目标任务相关性的特征,该通用评分规则实现了更精确的核近似。
  • 该算法在大规模学习场景中表现出鲁棒性和可扩展性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。