Skip to main content
QUICK REVIEW

[Paper Review] Transfer Learning for High-dimensional Linear Regression: Prediction, Estimation, and Minimax Optimality

Sai Li, Tommaso Cai|arXiv (Cornell University)|Jun 18, 2020
Statistical Methods and Inference32 references21 citations
TL;DR

This paper proposes a transfer learning framework for high-dimensional linear regression that improves prediction and estimation by leveraging informative auxiliary samples. It introduces Trans-Lasso, a data-driven method that adaptively aggregates estimators using sparsity in the contrast between target and auxiliary coefficients, achieving minimax optimal rates when informative sources are known and demonstrating robustness and improved performance in real gene expression data.

ABSTRACT

This paper considers the estimation and prediction of a high-dimensional linear regression in the setting of transfer learning, using samples from the target model as well as auxiliary samples from different but possibly related regression models. When the set of "informative" auxiliary samples is known, an estimator and a predictor are proposed and their optimality is established. The optimal rates of convergence for prediction and estimation are faster than the corresponding rates without using the auxiliary samples. This implies that knowledge from the informative auxiliary samples can be transferred to improve the learning performance of the target problem. In the case that the set of informative auxiliary samples is unknown, we propose a data-driven procedure for transfer learning, called Trans-Lasso, and reveal its robustness to non-informative auxiliary samples and its efficiency in knowledge transfer. The proposed procedures are demonstrated in numerical studies and are applied to a dataset concerning the associations among gene expressions. It is shown that Trans-Lasso leads to improved performance in gene expression prediction in a target tissue by incorporating the data from multiple different tissues as auxiliary samples.

Motivation & Objective

  • To develop a statistically optimal transfer learning method for high-dimensional linear regression using auxiliary samples from related but distinct models.
  • To characterize the similarity between the target model and auxiliary models via the sparsity of the difference in their coefficient vectors.
  • To establish minimax optimal rates of convergence for prediction and estimation when the set of informative auxiliary samples is known.
  • To design a data-driven procedure, Trans-Lasso, that adapts to unknown informative sets and remains robust to non-informative auxiliary samples.
  • To validate the method empirically through simulations and real-world application to gene expression data from the GTEx project.

Proposed method

  • Define the target model as a high-dimensional linear regression with sparse coefficients and use auxiliary samples from related models with potentially different coefficients.
  • Characterize informative auxiliary studies as those whose coefficient contrasts with the target are sparse in ℓq norm for q ∈ [0,1], formalizing similarity via sparsity of δ(k) = β − w(k).
  • Propose an Oracle Trans-Lasso estimator that uses known informative sets, achieving minimax optimal rates for prediction and estimation.
  • Develop Trans-Lasso as a data-driven aggregation procedure that combines candidate estimators from all auxiliary studies using a sparsity-based selection criterion.
  • Apply a penalized estimation framework with adaptive Lasso-type penalties to select and weight informative auxiliary samples based on their estimated similarity to the target.
  • Use cross-validation and theoretical analysis to demonstrate robustness and efficiency in the presence of non-informative auxiliary samples.

Experimental results

Research questions

  • RQ1Can transfer learning improve prediction and estimation accuracy in high-dimensional linear regression when auxiliary samples from related models are available?
  • RQ2What is the optimal rate of convergence for prediction and estimation in transfer learning for high-dimensional regression, and can it be achieved?
  • RQ3How can one adaptively identify and utilize informative auxiliary samples when the set of such samples is unknown?
  • RQ4To what extent does the proposed method remain robust to the inclusion of non-informative auxiliary samples?
  • RQ5Can the theoretical gains from transfer learning be empirically validated in real biological data, such as gene expression networks?

Key findings

  • The Oracle Trans-Lasso achieves minimax optimal rates for prediction and estimation when the set of informative auxiliary samples is known, with faster convergence rates than standard Lasso.
  • Trans-Lasso achieves a 17% average reduction in prediction error compared to the Lasso across multiple tissues in the GTEx data, demonstrating significant performance gains.
  • In tissues like Amygdala and Nucleus accumbens, Trans-Lasso shows substantial improvement, indicating effective knowledge transfer from related tissues.
  • In tissues like Pituitary, where the target model is less similar to others, the improvement is mild, confirming that transfer is limited to similar contexts.
  • Trans-Lasso outperforms the naive Trans-Lasso (which aggregates all auxiliary samples) in tissues like Frontal Cortex, where the naive version underperforms the Lasso, proving robustness to non-informative sources.
  • For 25 genes on Chromosome 21, Trans-Lasso achieves the best overall prediction performance, with significant gains in tissues including Cerebellar Hemisphere and Cortex.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.