[Paper Review] Distance Metrics for Measuring Joint Dependence with Application to Causal Inference
This paper introduces high-order distance covariance and joint distance covariance (JdCov) to measure joint dependence among d ≥ 2 random vectors, generalizing distance covariance to capture mutual independence through linear combinations of pairwise and higher-order dependence terms. The method enables consistent testing of joint independence in structural equation models via bootstrap-assisted inference, with asymptotically valid sampling distribution approximation and strong empirical power in simulation and real-data studies.
Many statistical applications require the quantification of joint dependence among more than two random vectors. In this work, we generalize the notion of distance covariance to quantify joint dependence among d >= 2 random vectors. We introduce the high order distance covariance to measure the so-called Lancaster interaction dependence. The joint distance covariance is then defined as a linear combination of pairwise distance covariances and their higher order counterparts which together completely characterize mutual independence. We further introduce some related concepts including the distance cumulant, distance characteristic function, and rank-based distance covariance. Empirical estimators are constructed based on certain Euclidean distances between sample elements. We study the large sample properties of the estimators and propose a bootstrap procedure to approximate their sampling distributions. The asymptotic validity of the bootstrap procedure is justified under both the null and alternative hypotheses. The new metrics are employed to perform model selection in causal inference, which is based on the joint independence testing of the residuals from the fitted structural equation models. The effectiveness of the method is illustrated via both simulated and real datasets.
Motivation & Objective
- To generalize distance covariance to quantify joint dependence among d ≥ 2 random vectors, beyond pairwise dependence.
- To develop a metric that fully characterizes mutual independence among d random vectors, including higher-order interaction effects.
- To provide a robust, nonparametric method for model selection in causal inference based on residual independence testing.
- To establish asymptotic theory and bootstrap procedures for valid inference on the proposed dependence metrics.
Proposed method
- Proposes high-order distance covariance to measure Lancaster interaction dependence among d random vectors.
- Defines joint distance covariance (JdCov) as a linear combination of pairwise and higher-order distance covariances, ensuring JdCov = 0 iff all d vectors are mutually independent.
- Constructs empirical estimators using U-statistics and V-statistics based on Euclidean distances between sample points.
- Introduces distance cumulant and distance characteristic function for equivalent characterization of independence.
- Develops a bootstrap procedure to approximate the sampling distribution of JdCov under both null and alternative hypotheses.
- Applies the method to test joint independence of residuals in structural equation models for causal discovery.
Experimental results
Research questions
- RQ1Can distance covariance be generalized to measure joint dependence among d ≥ 2 random vectors, including higher-order interaction effects?
- RQ2Does the proposed joint distance covariance (JdCov) fully characterize mutual independence among d random vectors?
- RQ3Is the bootstrap procedure asymptotically valid for approximating the sampling distribution of JdCov under both the null and alternative hypotheses?
- RQ4How effective is JdCov in detecting joint dependence in high-dimensional settings and in model selection for causal structure learning?
Key findings
- The proposed joint distance covariance (JdCov) is zero if and only if the d random vectors are mutually independent, ensuring complete characterization of joint independence.
- In simulations with n = 100 and d = 10, the bootstrap-assisted test achieved 100% power for detecting dependence in Example 5.2 and 98.5% power in Example 5.1.
- For d = 3 and n = 50, the method achieved 98.5% empirical power in Example 5.3, demonstrating strong detection capability even in low sample sizes.
- The bootstrap procedure demonstrated asymptotic validity, with empirical sizes close to nominal levels (e.g., 10.3% and 5.3% for 10% and 5% nominal levels in Table 6).
- The method showed high power in detecting complex dependence structures, including non-monotonic and higher-order interactions, outperforming classical rank-based measures.
- In Example 5.4 with d = 5 and n = 200, the test maintained 80.4% power for detecting dependence with a 10% significance level, confirming robustness in larger dimensions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.