[Paper Review] Deriving the Variance of the Discrete Fourier Transform Test Using Parseval's Theorem
This paper derives the theoretical variance of the Discrete Fourier Transform Test (DFTT) statistic in NIST SP800-22 using Parseval’s theorem under specific assumptions. It shows that the variance of the count $ N_1 $ of DFT magnitudes below threshold $ \sqrt{-n\log(0.05)} $ is approximately $ (0.95)(0.05)n/a $, with $ a \approx 3.7903 $, resolving discrepancies in prior empirical estimates and validating the result through extensive numerical experiments.
The discrete Fourier transform test is a randomness test included in NIST SP800-22. However, the variance of the test statistic is smaller than expected and the theoretical value of the variance is not known. Hitherto, the mechanism explaining why the former variance is smaller than expected has been qualitatively explained based on Parseval's theorem. In this paper, we explore this quantitatively and derive the variance using Parseval's theorem under particular assumptions. Numerical experiments are then used to show that this derived variance is robust.
Motivation & Objective
- To resolve the long-standing issue of the unknown theoretical variance of the DFTT test statistic, which is observed to be significantly smaller than the expected $ (0.95)(0.05)n/2 $.
- To provide a theoretically grounded explanation for the reduced variance using Parseval’s theorem, moving beyond qualitative insights.
- To derive a precise theoretical value for the variance scaling factor $ a $ in $ \text{Var}[N_1] = (0.95)(0.05)n/a $, under plausible assumptions about the distribution of DFT magnitudes.
- To validate the derived variance through large-scale numerical experiments using the Mersenne Twister random number generator.
Proposed method
- Assumes that the DFT magnitudes $ |f_j| $ for $ j = 1, \dots, m-1 $ follow a distribution approximated by a scaled chi-squared distribution with 2 degrees of freedom.
- Applies Parseval’s theorem to relate the sum of squared DFT magnitudes to the energy of the input sequence, establishing constraints on the distribution of $ |f_j| $ values.
- Models the indicator variables $ F_j = \mathbf{1}_{\{|f_j| \leq T\}} $ for $ T = \sqrt{-n\log(0.05)} $, and computes the variance of $ N_1 = \sum_{j=0}^{m-1} F_j $ by decomposing it into variances and covariances.
- Derives analytical expressions for $ \text{Var}[F_j] $ and $ \text{Corr}[F_i, F_j] $ under the assumed distribution, leading to a closed-form expression for $ \text{Var}[N_1] $.
- Solves for the scaling factor $ a $ such that $ \text{Var}[N_1] = (0.95)(0.05)n/a $, yielding $ a \approx 3.7903 $ as $ m \to \infty $.
- Validates the theoretical result via numerical experiments using $ 10^8 $ sequences of length $ 10^6 $, computing empirical correlation and variance to estimate $ a $.
Experimental results
Research questions
- RQ1What is the theoretical value of the variance of the DFTT test statistic $ N_1 $, given that prior estimates were based on empirical observations only?
- RQ2How does Parseval’s theorem explain the observed reduction in variance compared to the nominal $ (0.95)(0.05)n/2 $?
- RQ3Can a closed-form analytical expression for the variance of $ N_1 $ be derived under reasonable assumptions about the distribution of DFT magnitudes?
- RQ4Is the derived variance value robust and consistent with empirical data from high-quality random number generators?
- RQ5Does the derived scaling factor $ a \approx 3.7903 $ provide a better estimate than previous empirical values such as 3.7879 or 3.8?
Key findings
- The theoretical variance of $ N_1 $, the count of DFT magnitudes below threshold $ \sqrt{-n\log(0.05)} $, is derived as $ \text{Var}[N_1] = (0.95)(0.05)n/a $, with $ a \approx 3.7903 $, based on Parseval’s theorem and distributional assumptions.
- The derived value of $ a \approx 3.7903 $ is consistent with numerical experiments using $ 10^8 $ sequences generated by the Mersenne Twister, yielding an empirical mean of $ a = 3.790217 \pm 0.000527 $.
- The correlation between indicator variables $ F_i $ and $ F_j $ for different DFT frequencies is non-zero and contributes significantly to the reduced variance, explaining the discrepancy from the nominal $ (0.95)(0.05)n/2 $.
- The leading-order terms of the correlation coefficient $ C[F_i, F_j] $ derived from the theoretical model match those observed in simulations, validating the model’s accuracy.
- The derived variance is more accurate than prior empirical estimates such as $ a \approx 3.7879 $ or $ a \approx 3.8 $, and is recommended for use in DFTT to improve the precision of randomness testing.
- The result supports the use of $ d = \frac{N_1 - 0.95n/2}{\sqrt{(0.95)(0.05)n/3.7903}} $ in the test statistic to ensure proper standardization under the null hypothesis of randomness.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.