Skip to main content
QUICK REVIEW

[Paper Review] Asymptotic minimaxity of False Discovery Rate thresholding for sparse exponential data

David L. Donoho, Jiashun Jin|arXiv (Cornell University)|Feb 14, 2006
Statistical Methods and InferenceMathematics16 references37 citations
TL;DR

This paper establishes the asymptotic minimaxity of False Discovery Rate (FDR) thresholding for estimating sparse exponential means, showing that FDR control with q ≤ 1/2 yields nearly optimal estimation performance in terms of mean-squared error on the log-scale. The method adaptively thresholds data to recover signals in high-dimensional, sparse exponential models, with risk approaching the minimax lower bound as sparsity increases (η → 0).

ABSTRACT

We apply FDR thresholding to a non-Gaussian vector whose coordinates X_i, i=1,..., n, are independent exponential with individual means $\mu_i$. The vector $\mu =(\mu_i)$ is thought to be sparse, with most coordinates 1 but a small fraction significantly larger than 1; roughly, most coordinates are simply `noise,' but a small fraction contain `signal.' We measure risk by per-coordinate mean-squared error in recovering $\log(\mu_i)$, and study minimax estimation over parameter spaces defined by constraints on the per-coordinate p-norm of $\log(\mu_i)$: $\frac{1}{n}\sum_{i=1}^n\log^p(\mu_i)\leq \eta^p$. We show for large n and small $\eta$ that FDR thresholding can be nearly Minimax. The FDR control parameter 0<q<1 plays an important role: when $q\leq 1/2$, the FDR estimator is nearly minimax, while choosing a fixed q>1/2 prevents near minimaxity. These conclusions mirror those found in the Gaussian case in Abramovich et al. [Ann. Statist. 34 (2006) 584--653]. The techniques developed here seem applicable to a wide range of other distributional assumptions, other loss measures and non-i.i.d. dependency structures.

Motivation & Objective

  • To study minimax estimation of sparse exponential means where most µi = 1 but a few are significantly larger.
  • To evaluate the performance of FDR thresholding as a data-adaptive estimation rule under ℓp-constrained sparsity.
  • To determine under what conditions FDR thresholding achieves asymptotic minimax risk in non-Gaussian, sparse exponential models.
  • To generalize results from Gaussian settings to exponential noise, focusing on log-scale loss and thresholding rules.
  • To establish conditions under which FDR thresholding is asymptotically minimax, particularly the role of the FDR control parameter q.

Proposed method

  • Models observations Xi ∼ Exp(µi) with sparse µi, where most µi = 1 and a small fraction are >1, constrained by ℓp-norm on log(µi).
  • Measures estimation risk via per-coordinate mean-squared error on the log-scale: (1/n)∑(log ˆµi − log µi)².
  • Defines minimax risk R∗n as the infimum over all estimators of the worst-case risk over ℓp-balls of radius η.
  • Proposes FDR thresholding estimator ˆµFDR,q,n using data-dependent threshold tFDR = X(kFDR), where kFDR is the largest k with X(k) ≥ −log(qk/n).
  • Applies Simes' procedure to control FDR at level q, ensuring expected proportion of false discoveries ≤ q.
  • Analyzes asymptotic behavior as n → ∞ followed by η → 0, using extreme value theory and empirical process tools on order statistics.

Experimental results

Research questions

  • RQ1Under what conditions is FDR thresholding asymptotically minimax for sparse exponential data?
  • RQ2How does the choice of FDR control parameter q affect the minimax risk of the FDR estimator?
  • RQ3Can FDR thresholding achieve near-optimal estimation performance in non-Gaussian, sparse models with heavy-tailed or non-normal noise?
  • RQ4What is the role of the sparsity parameter η and the ℓp-norm constraint in determining the minimax risk and estimator performance?
  • RQ5How does FDR thresholding compare to oracle thresholding and other adaptive estimation rules in high-dimensional sparse exponential models?

Key findings

  • When 0 < q ≤ 1/2, the FDR estimator ˆµFDR,q,n is asymptotically minimax: the ratio of its worst-case risk to the minimax risk tends to 1 as n → ∞ and η → 0.
  • When q > 1/2, the FDR estimator is not asymptotically minimax, and the risk ratio converges to q/(1−q) > 1, indicating a strict performance gap.
  • The minimax risk R∗n(Mn,p(η)) behaves asymptotically as ηp log²⁻ᵖ log(1/η), establishing the sharp rate of convergence.
  • The optimal threshold t0(p,η) for oracle thresholding scales as p log(1/η) + p log log(1/η) · (1 + o(1)) as η → 0.
  • FDR thresholding performs well even in finite samples, with empirical risk close to η log log(1/η) for q < 1/2 and increasing only moderately for q near 1/2.
  • The method extends to other non-Gaussian models, including sparse Poisson means and additive noise with Gumbel or double-exponential distributions, provided tails are light enough.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.