Skip to main content
QUICK REVIEW

[Paper Review] Performance of information criteria used for model selection of Hawkes process models of financial data

J. M. Chen, Alan G. Hawkes|arXiv (Cornell University)|Feb 20, 2017
Point processes and geometric inequalities4 citations
TL;DR

This study evaluates the performance of AIC, BIC, and HQ information criteria for selecting the order of Hawkes processes with mixed exponential kernels in high-frequency financial data. Using Monte Carlo simulations with realistic parameter sets and sample sizes, it finds that AIC outperforms BIC and HQ in small samples, while consistent criteria (BIC, HQ) achieve near-100% correct model selection in large samples, though AIC shows a non-monotonic success rate with increasing sample size.

ABSTRACT

We test three common information criteria (IC) for selecting the order of a Hawkes process with an intensity kernel that can be expressed as a mixture of exponential terms. These processes find application in high-frequency financial data modelling. The information criteria are Akaike's information criterion (AIC), the Bayesian information criterion (BIC) and the Hannan-Quinn criterion (HQ). Since we work with simulated data, we are able to measure the performance of model selection by the success rate of the IC in selecting the model that was used to generate the data. In particular, we are interested in the relation between correct model selection and underlying sample size. The analysis includes realistic sample sizes and parameter sets from recent literature where parameters were estimated using empirical financial intra-day data. We compare our results to theoretical predictions and similar empirical findings on the asymptotic distribution of model selection for consistent and inconsistent IC.

Motivation & Objective

  • To assess the performance of AIC, BIC, and HQ in selecting the correct model order for Hawkes processes with mixed exponential intensity kernels.
  • To investigate how model selection accuracy varies with sample size in realistic financial data settings.
  • To compare empirical results against theoretical predictions on consistency and asymptotic behavior of information criteria.
  • To evaluate the robustness of information criteria under varying parameter configurations and sample sizes drawn from real intra-day financial data.
  • To examine the potential for non-monotonic behavior in model selection success rates, particularly for AIC.

Proposed method

  • Simulated Hawkes processes with intensity kernels as weighted sums of up to three exponential terms using the thinning algorithm.
  • Estimated model parameters via maximum likelihood estimation (MLE) using the fmincon optimization routine in MATLAB.
  • Calculated AIC, BIC, and HQ values for each candidate model order using the log-likelihood from MLE.
  • Conducted Monte Carlo experiments with repeated simulations across multiple sample sizes (T = 500 to T = 21,600) and two realistic parameter sets.
  • Measured model selection success rate as the proportion of simulations where the IC selected the true data-generating model order.
  • Evaluated root mean squared error (RMSE) of parameter estimates across simulations to assess estimation accuracy.

Experimental results

Research questions

  • RQ1How do AIC, BIC, and HQ perform in selecting the correct model order for Hawkes processes with mixed exponential kernels in small to moderate sample sizes?
  • RQ2Does the success rate of model selection for AIC, BIC, and HQ increase monotonically with sample size, or are there non-monotonic patterns?
  • RQ3How does the performance of information criteria compare to theoretical expectations regarding consistency and asymptotic behavior?
  • RQ4To what extent do parameter estimation errors (RMSE) affect model selection accuracy, especially for higher-order components?
  • RQ5Is the combined AICc/AIC rule more effective than standalone AIC in small-sample settings?

Key findings

  • For small samples (T = 500), AIC achieved the highest success rate in model selection, outperforming BIC and HQ, consistent with its inconsistency and bias toward overfitting.
  • BIC and HQ showed monotonic improvement in model selection success rate with increasing sample size, reaching nearly 100% correct selection at T = 2000.
  • AIC exhibited a non-monotonic success rate pattern, with a dip in performance around T = 1800, suggesting potential overfitting tendencies in intermediate sample sizes.
  • For larger samples (T ≥ 3600), consistent IC (BIC and HQ) achieved over 90% correct model selection, with BIC and HQ both approaching 100% at T = 21,600.
  • The success rate of AIC decreased slightly to about 94% at T = 21,600, with a 6% probability of overestimating the model order, indicating a tendency to select overly complex models.
  • Parameter estimation RMSE was high for smaller samples (T = 600, 900), especially for the second exponential component, but decreased significantly as sample size increased, improving model selection accuracy.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.