Skip to main content
QUICK REVIEW

[Paper Review] Associating High-Dimensional Longitudinal Datasets through an Efficient Cross-Covariance Decomposition

Jianbin Tan, Pixu Shi|arXiv (Cornell University)|Jan 19, 2026
Single-cell and spatial transcriptomics0 citations
TL;DR

FACD is a framework for high-dimensional longitudinal data that learns time-varying cross-covariances via data-adaptive bases and SVD, with sparsity for feature selection and theoretical guarantees; it outperforms related methods in simulations and reveals dynamic cross-omic associations in a longitudinal study.

ABSTRACT

Understanding associations between paired high-dimensional longitudinal datasets is a fundamental yet challenging problem that arises across scientific domains, including longitudinal multi-omic studies. The difficulty stems from the complex, time-varying cross-covariance structure coupled with high dimensionality, which complicates both model formulation and statistical estimation. To address these challenges, we propose a new framework, termed Functional-Aggregated Cross-covariance Decomposition (FACD), tailored for canonical cross-covariance analysis between paired high-dimensional longitudinal datasets through a statistically efficient and theoretically grounded procedure. Unlike existing methods that are often limited to low-dimensional data or rely on explicit parametric modeling of temporal dynamics, FACD adaptively learns temporal structure by aggregating signals across features and naturally accommodates variable selection to identify the most relevant features associated across datasets. We establish statistical guarantees for FACD and demonstrate its advantages over existing approaches through extensive simulation studies. Finally, we apply FACD to a longitudinal multi-omic human study, revealing blood molecules with time-varying associations across omic layers during acute exercise.

Motivation & Objective

  • Motivate the need to understand dynamic associations between paired high-dimensional longitudinal datasets.
  • Introduce FACD as a statistically efficient framework that learns time-varying cross-covariance through adaptive basis expansion.
  • Incorporate sparsity to identify key features linking the two datasets.
  • Provide theoretical guarantees for FACD and validate its performance via simulations.
  • Demonstrate the method on a longitudinal multi-omic study to uncover time-varying cross-omic associations.

Proposed method

  • Formulate canonical cross-covariance analysis for high-dimensional functional data as a spectral decomposition of the cross-covariance operator.
  • Represent the cross-covariance via functional basis expansion using data-adaptive bases rather than pre-specified bases.
  • Construct data-adaptive bases through extraction from kernels H_X and H_Y derived from R_XY, enabling a tractable matrix SVD instead of high-dimensional operator decomposition.
  • Apply a truncation approach to approximate R_XY with a finite set of basis functions and solve a matrix SVD on the resulting Gamma matrix.
  • Introduce sparsity through a penalized optimization (grouped lasso-like penalties) to select nonzero feature loadings while constraining the loadings to unit norms.
  • Handle irregular and sparse time observations by using spline-based estimation of mean functions and cross-covariance kernels, with smoothing penalties and GCV for tuning.

Experimental results

Research questions

  • RQ1How can we robustly estimate canonical cross-covariance components between two high-dimensional longitudinal datasets without relying on restrictive parametric temporal models?
  • RQ2Can data-adaptive bases and a tractable SVD-based decomposition accurately recover time-varying cross-dataset associations in high dimensions?
  • RQ3Does incorporating sparsity improve interpretability and identification of key shared features across datasets?
  • RQ4What theoretical guarantees can be established for FACD in terms of existence, stability, and approximation error?
  • RQ5How does FACD perform relative to existing methods in simulations and real longitudinal multi-omic data?

Key findings

  • FACD provides a theoretically grounded framework that reduces infinite-dimensional operator decomposition to a finite SVD problem.
  • The cross-covariance between high-dimensional longitudinal data can be represented via data-adaptive bases extracted from kernels H_X and H_Y.
  • A truncation-based approximation coupled with SVD yields computable canonical loadings and scores.
  • Sparsity constraints enable identification of a small subset of features driving cross-dataset associations.
  • Simulation studies show FACD has advantages over related methods in recovering canonical components.
  • Applied to a longitudinal multi-omic study, FACD reveals time-varying associations across omic layers during acute exercise.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.