[Paper Review] Granger Causality Networks for Categorical Time Series
This paper proposes a convex reformulation of the Mixture Transition Distribution (MTD) model for learning Granger causality networks in multivariate categorical time series, overcoming its traditional non-convexity and identifiability issues. The method enables efficient, high-dimensional estimation via a projected gradient algorithm with Dykstra-style projections, outperforming the multinomial logistic transition distribution (mLTD) in interpretability and sparsity while maintaining strong performance under model misspecification.
We present a new framework for learning Granger causality networks for multivariate categorical time series, based on the mixture transition distribution (MTD) model. Traditionally, MTD is plagued by a nonconvex objective, non-identifiability, and presence of many local optima. To circumvent these problems, we recast inference in the MTD as a convex problem. The new formulation facilitates the application of MTD to high-dimensional multivariate time series. As a baseline, we also formulate a multi-output logistic autoregressive model (mLTD), which while a straightforward extension of autoregressive Bernoulli generalized linear models, has not been previously applied to the analysis of multivariate categorial time series. We develop novel identifiability conditions of the MTD model and compare them to those for mLTD. We further devise novel and efficient optimization algorithm for the MTD based on the new convex formulation, and compare the MTD and mLTD in both simulated and real data experiments. Our approach simultaneously provides a comparison of methods for network inference in categorical time series and opens the door to modern, regularized inference with the MTD model.
Motivation & Objective
- To address the long-standing computational and identifiability challenges of the Mixture Transition Distribution (MTD) model for multivariate categorical time series.
- To develop a convex optimization framework for MTD that enables regularized, high-dimensional inference of Granger causality networks.
- To establish novel identifiability conditions for MTD and compare them with those of the multinomial logistic transition distribution (mLTD).
- To provide a computationally efficient and interpretable alternative to existing models for categorical time series network inference.
- To enable scalable application of MTD to modern, high-dimensional datasets through a novel reparameterization and optimization strategy.
Proposed method
- Reformulate the MTD model as a convex optimization problem via a novel reparameterization that transforms non-convex constraints into convex ones.
- Introduce a penalized likelihood framework with sparsity-inducing penalties to promote sparse Granger causality networks.
- Develop a projected gradient algorithm using Dykstra's splitting method to efficiently project onto the convex constraint set of the MTD model.
- Use convex penalties to enforce identifiability constraints, replacing the previous non-convex and hard-to-satisfy conditions.
- Compare the MTD approach with a novel baseline: the multi-output logistic transition distribution (mLTD), which extends autoregressive GLMs to multinomial outcomes.
- Apply the MTD and mLTD models to simulated data and real-world datasets, including Bach chorale harmony, to evaluate performance and interpretability.
Experimental results
Research questions
- RQ1Can the non-convex MTD model be reformulated as a convex optimization problem to enable scalable, high-dimensional inference?
- RQ2What are the new identifiability conditions for the MTD model, and how do they compare to those of the mLTD model?
- RQ3How does the performance of the convex MTD framework compare to the mLTD model in terms of network sparsity, interpretability, and robustness under model misspecification?
- RQ4Can the proposed projected gradient algorithm with Dykstra’s method efficiently handle high-dimensional categorical time series data?
- RQ5To what extent can the MTD model recover true causal structures in multivariate categorical time series, especially when the true data-generating process is misspecified?
Key findings
- The proposed convex MTD formulation enables efficient, scalable inference in high-dimensional multivariate categorical time series, overcoming the prior limitations of non-convex optimization and local optima.
- The new identifiability conditions for MTD are derived and shown to be enforceable via convex penalties, ensuring model consistency.
- The projected gradient algorithm with Dykstra’s splitting method achieves fast and accurate projection onto the MTD constraint set, enabling large-scale applications.
- In real data analysis of Bach chorales, the MTD model recovered a sparse, interpretable network with strong connections between harmony notes and chord, consistent with music theory.
- The MTD model produced a sparser and more interpretable network than mLTD, which suffered from non-intuitive link strength metrics due to the nonlinearity of the softmax function.
- Simulations and real data show that both MTD and mLTD perform well under model misspecification, but MTD offers superior interpretability and structure recovery.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.