Skip to main content
QUICK REVIEW

[Paper Review] High-dimensional nonparametric density estimation via symmetry and shape constraints

Min Xu, Richard J. Samworth|arXiv (Cornell University)|Mar 14, 2019
Point processes and geometric inequalities40 references8 citations
TL;DR

This paper proposes a high-dimensional nonparametric density estimation method using symmetry and shape constraints—specifically, K-homothetic log-concave densities—where super-level sets are scalar multiples of a convex body K. By leveraging log-concavity and homotheticity, the method achieves a worst-case squared Hellinger risk of O(n⁻⁴/⁵) independent of dimension p, evading the curse of dimensionality, and adapts to nearly parametric rates when the generator function is piecewise linear with few segments.

ABSTRACT

We tackle the problem of high-dimensional nonparametric density estimation by taking the class of log-concave densities on $\mathbb{R}^p$ and incorporating within it symmetry assumptions, which facilitate scalable estimation algorithms and can mitigate the curse of dimensionality. Our main symmetry assumption is that the super-level sets of the density are $K$-homothetic (i.e. scalar multiples of a convex body $K \subseteq \mathbb{R}^p$). When $K$ is known, we prove that the $K$-homothetic log-concave maximum likelihood estimator based on $n$ independent observations from such a density has a worst-case risk bound with respect to, e.g., squared Hellinger loss, of $O(n^{-4/5})$, independent of $p$. Moreover, we show that the estimator is adaptive in the sense that if the data generating density admits a special form, then a nearly parametric rate may be attained. We also provide worst-case and adaptive risk bounds in cases where $K$ is only known up to a positive definite transformation, and where it is completely unknown and must be estimated nonparametrically. Our estimation algorithms are fast even when $n$ and $p$ are on the order of hundreds of thousands, and we illustrate the strong finite-sample performance of our methods on simulated data.

Motivation & Objective

  • Address the curse of dimensionality in high-dimensional nonparametric density estimation by incorporating shape and symmetry constraints.
  • Develop scalable, tuning-free estimation algorithms for high-dimensional data with n and p in the hundreds of thousands.
  • Establish theoretical risk bounds under various assumptions about the unknown super-level set K and centering vector µ.
  • Demonstrate adaptation to low-complexity densities (e.g., piecewise linear generator functions) with nearly parametric rates.
  • Provide a computationally efficient plug-in approach for estimating K and µ when they are unknown, including a novel convex-hull-based algorithm for nonparametric K estimation.

Proposed method

  • Define the class of K-homothetic log-concave densities as f(x) = exp(φ(‖x−µ‖_K)) where φ is concave and decreasing, and ‖·‖_K is the Minkowski functional of a convex body K.
  • Use the maximum likelihood estimator (MLE) within the K-homothetic log-concave class, which is computationally tractable due to the absence of tuning parameters.
  • Establish a homothetic, log-concave projection ψ*_{K,µ} from the class of distributions with finite mean to the K-homothetic log-concave class, enabling the plug-in estimator ˆfn = ψ*_{K̂,μ̂}(P_n).
  • Propose a nonparametric estimator of K by computing the convex hull of boundary estimates in random directions, where boundary estimates are the average Euclidean norm of data points in cones around each direction.
  • Derive risk bounds using the divergence measure d²_X(ˆfn, f₀) = (1/n)∑ log(ˆfn(X_i)/f₀(X_i)), which upper bounds Kullback–Leibler, squared Hellinger, and total variation risks.
  • Apply geometric probability tools, including Hausdorff distance and scale distortion (dscale), to control the deviation between estimated and true super-level sets K.

Experimental results

Research questions

  • RQ1Can symmetry and shape constraints be used to achieve faster convergence rates in high-dimensional nonparametric density estimation?
  • RQ2What is the theoretical risk bound for the K-homothetic log-concave MLE when K and µ are known, and does it depend on dimension p?
  • RQ3How does the estimator adapt to densities with simple generator functions (e.g., piecewise linear φ)?
  • RQ4What are the risk bounds when K and µ are unknown and must be estimated, particularly in semiparametric and nonparametric settings?
  • RQ5Can a scalable, tuning-free algorithm be constructed for high-dimensional density estimation under homothetic and log-concave constraints?

Key findings

  • When K and µ are known, the K-homothetic log-concave MLE achieves a worst-case squared Hellinger risk of O(n⁻⁴/⁵), independent of dimension p.
  • When the true density corresponds to a piecewise linear generator function with k segments, the risk is bounded by O(k/n log⁵/⁴(en/k)), approaching a nearly parametric rate for small k.
  • In the semiparametric setting where K = Σ₀¹ᐟ²K₀ for known K₀ and unknown Σ₀, the worst-case squared Hellinger risk is O(p³ᐟ²/n¹ᐟ²) up to polylogarithmic factors.
  • For smooth or 1-affine generator functions in the semiparametric setting, the adaptation rates are O(n⁻⁴ᐟ⁵ + p³/n) and O(p³/n), respectively, up to logarithmic factors.
  • In the nonparametric setting where K is arbitrary, the proposed convex-hull-based estimator of K achieves a worst-case squared Hellinger risk of O((log M / M)¹ᐟᵖ⁻¹) when M random directions are used.
  • Empirical studies confirm strong finite-sample performance, with fast computation even for n, p ~ 10⁵, and the method outperforms standard nonparametric methods in high dimensions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.