Skip to main content
QUICK REVIEW

[论文解读] Multitask Learning using Task Clustering with Applications to Predictive Modeling and GWAS of Plant Varieties

Ming Yu, Addie Thompson|arXiv (Cornell University)|Oct 4, 2017
Spectroscopy and Chemometric Analyses参考文献 16被引用 4
一句话总结

该论文提出了一种新颖的多任务学习框架,通过凸聚类正则化联合学习稀疏线性模型与任务之间的层次树结构。该方法能从数据中自动推断任务关系,在植物性状预测和基于遥感数据的全基因组关联研究(GWAS)中提升预测准确性,并揭示具有生物学意义的分组关系。

ABSTRACT

Inferring predictive maps between multiple input and multiple output variables or tasks has innumerable applications in data science. Multi-task learning attempts to learn the maps to several output tasks simultaneously with information sharing between them. We propose a novel multi-task learning framework for sparse linear regression, where a full task hierarchy is automatically inferred from the data, with the assumption that the task parameters follow a hierarchical tree structure. The leaves of the tree are the parameters for individual tasks, and the root is the global model that approximates all the tasks. We apply the proposed approach to develop and evaluate: (a) predictive models of plant traits using large-scale and automated remote sensing data, and (b) GWAS methodologies mapping such derived phenotypes in lieu of hand-measured traits. We demonstrate the superior performance of our approach compared to other methods, as well as the usefulness of discovering hierarchical groupings between tasks. Our results suggest that richer genetic mapping can indeed be obtained from the remote sensing data. In addition, our discovered groupings reveal interesting insights from a plant science perspective.

研究动机与目标

  • 开发一种监督式多任务学习方法,能够以端到端方式同时估计任务特定参数并从数据中推断层次化任务结构。
  • 通过利用层次聚类中的任务相关性,提升大规模遥感数据下植物性状预测建模的性能。
  • 通过使用从遥感数据中提取的衍生表型而非人工测量,实现更强大的全基因组关联研究(GWAS)。
  • 提供一种统计上稳健、凸优化的框架,具备收敛性保证和可解释的任务分组。
  • 在合成数据和真实世界高通量表型应用中,证明该方法优于现有方法。

提出的方法

  • 该方法采用凸优化框架,将稀疏线性回归与任务参数上的凸聚类惩罚相结合,以诱导出层次树结构。
  • 其目标函数通过正则化形式,同时鼓励任务特定系数的稀疏性以及通过连续正则化路径实现相似任务的分组。
  • 该方法采用邻近分解算法求解所得的凸优化问题,具有已证明的数值收敛性。
  • 层次结构由数据自动学习得到,树中的每个节点代表一个共享参数组,根节点代表全局模型。
  • 该方法使用由调优参数 λ₂ 索引的连续解路径,通过在不同阈值处截断树结构,可获得不同的任务分组。
  • 该框架被应用于植物性状的预测建模及下游GWAS分析,使用从遥感数据中提取的基于直方图的表型。

实验结果

研究问题

  • RQ1能否在完全监督、数据驱动的方式下,使多任务学习框架同时学习任务参数与层次化任务结构?
  • RQ2与现有多任务学习方法相比,该方法在植物性状建模中的预测准确性如何提升?
  • RQ3自动发现的任务分组能否揭示植物性状或基因组区域之间的生物学意义关系?
  • RQ4使用遥感数据中提取的衍生表型,是否能比传统人工测量的性状实现更强大的遗传定位?
  • RQ5该方法在收敛性、可扩展性和鲁棒性方面的统计与计算性能如何?

主要发现

  • 所提方法在预测性能上表现更优,在合成数据上相比基线方法(包括不学习任务结构的方法)具有更低的RMSE。
  • 该方法成功地将遥感直方图区间聚类为具有生物学可解释性的组别,且高丰度区间在层次结构中后期才合并,反映出更强的信号。
  • 在第7号染色体上,该方法识别出与已知的Dwarf3基因共定位的SNPs,验证了其生物学相关性。
  • 在第9号染色体上,该方法发现了潜在与冠层闭合、叶片分布和开花相关的新型候选区域,提示新的遗传见解。
  • 在可扩展性和稳定性方面,该方法优于Kang等人提出的方法,尤其在高维设置下,后者无法收敛。
  • 所发现的任务分组揭示了有意义的植物科学洞见,例如与植物结构和生长模式相关的区间被聚类在一起。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。