Skip to main content
QUICK REVIEW

[Paper Review] A new method for augmenting short time series, with application to pain events in sickle cell disease

Kumar Utkarsh, Nirmish Shah|arXiv (Cornell University)|Jan 8, 2026
Hemoglobinopathies and Related Disorders0 citations
TL;DR

The paper presents a data augmentation framework that pools statistically similar sparse time series to improve Hawkes vs Poisson model discrimination and parameter estimation, applied to sickle cell disease pain event data.

ABSTRACT

Researchers across different fields, including but not limited to ecology, biology, and healthcare, often face the challenge of sparse data. Such sparsity can lead to uncertainties, estimation difficulties, and potential biases in modeling. Here we introduce a novel data augmentation method that combines multiple sparse time series datasets when they share similar statistical properties, thereby improving parameter estimation and model selection reliability. We demonstrate the effectiveness of this approach through validation studies comparing Hawkes and Poisson processes, followed by application to subjective pain dynamics in patients with sickle cell disease (SCD), a condition affecting millions worldwide, particularly those of African, Mediterranean, Middle Eastern, and Indian descent.

Motivation & Objective

  • Address the challenge of sparse time-series data hindering reliable model fitting and selection.
  • Develop a method to identify statistically similar datasets and pool them to form an augmented likelihood.
  • Validate the approach with simulations differentiating Hawkes from Poisson processes.
  • Apply the method to real-world sickle cell disease pain-event data to uncover temporal dynamics.

Proposed method

  • Model a self-exciting Hawkes process with exponential memory kernel and a compensatory term for unobserved past events (Eq. 2).
  • Compare Hawkes versus Poisson models using maximum likelihood and Akaike Information Criterion (AIC) for model selection.
  • Use a two-sample Kolmogorov-Smirnov (KS) test on interarrival times to identify datasets with similar distributions.
  • Define a collective likelihood that multiplies individual likelihoods across statistically similar datasets (Eq. 5).
  • Apply the augmentation workflow to sparse datasets, then re-estimate parameters and re-evaluate model support.
Figure 1: Visual guide to shifted Hawkes process parameters and intensity dynamics. Characterization of the parameters introduced in Eq. ( 2 ) (see also Table 1 ). The peaks represent event arrivals in real-time. The shaded area represents the history not captured in the observed data. In this examp
Figure 1: Visual guide to shifted Hawkes process parameters and intensity dynamics. Characterization of the parameters introduced in Eq. ( 2 ) (see also Table 1 ). The peaks represent event arrivals in real-time. The shaded area represents the history not captured in the observed data. In this examp

Experimental results

Research questions

  • RQ1Can sparse time-series data be augmented by pooling statistically similar datasets to improve model discrimination between Hawkes and Poisson processes?
  • RQ2Does the proposed augmentation approach improve parameter estimation for Hawkes model parameters (lambda_0, alpha, delta) under sparsity?
  • RQ3How does the augmentation affect model selection in real-world SCD pain-event data compared to single-series analysis?
  • RQ4What are the limitations and conditions under which KS-based similarity grouping yields reliable collective likelihood inferences?

Key findings

  • Augmented datasets shift model selection from inconclusive or Poisson-favored toward Hawkes-favored, increasing confidence (>95%) in many cases.
  • Parameter estimates from augmented data recover Hawkes parameters comparably to equivalent-length continuous data, improving robustness under sparsity.
  • In simulations, augmentation moves results outside the inconclusive region in Delta AIC for both Poisson and Hawkes processes.
  • Applied to 39 SCD patients, augmented fits show Hawkes preference in 36 of 39 cases, versus 28/39 for single-series fits.
  • The memory timescale delta^{-1} observed in real data ranges from 30 seconds to 6 minutes, informing risk-period duration after pain events.
Figure 2: Minimum dataset size required for reliable Hawkes vs. Poisson model discrimination. Number of data points needed to distinguish Hawkes model from Poisson. The black dashed line is for basic preference ( $\mathcal{L}=1$ ), whereas the green dashed line is for 95% confidence ( $\mathcal{L}=0
Figure 2: Minimum dataset size required for reliable Hawkes vs. Poisson model discrimination. Number of data points needed to distinguish Hawkes model from Poisson. The black dashed line is for basic preference ( $\mathcal{L}=1$ ), whereas the green dashed line is for 95% confidence ( $\mathcal{L}=0

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.