Skip to main content
QUICK REVIEW

[Paper Review] CompEngine: a self-organizing, living library of time-series data

Ben Fulcher, Carl H Lubba|arXiv (Cornell University)|May 3, 2019
Time Series Analysis and Forecasting15 references4 citations
TL;DR

CompEngine is a self-organizing, web-based library that uses a canonical feature-based representation to embed diverse time-series data—empirical, synthetic, and interdisciplinary—into a shared, low-dimensional space, enabling automatic discovery of meaningful connections across disciplines. The platform dynamically links similar time series in real time, incentivizing data sharing and fostering interdisciplinary collaboration by revealing unexpected analogies between real-world systems and model simulations.

ABSTRACT

Modern biomedical applications often involve time-series data, from high-throughput phenotyping of model organisms, through to individual disease diagnosis and treatment using biomedical data streams. Data and tools for time-series analysis are developed and applied across the sciences and in industry, but meaningful cross-disciplinary interactions are limited by the challenge of identifying fruitful connections. Here we introduce the web platform, CompEngine, a self-organizing, living library of time-series data that lowers the barrier to forming meaningful interdisciplinary connections between time series. Using a canonical feature-based representation, CompEngine places all time series in a common space, regardless of their origin, allowing users to upload their data and immediately explore interdisciplinary connections to other data with similar properties, and be alerted when similar data is uploaded in the future. In contrast to conventional databases, which are organized by assigned metadata, CompEngine incentivizes data sharing by automatically connecting experimental and theoretical scientists across disciplines based on the empirical structure of their data. CompEngine's growing library of interdisciplinary time-series data also facilitates comprehensively characterization of algorithm performance across diverse types of data, and can be used to empirically motivate the development of new time-series analysis algorithms.

Motivation & Objective

  • To address the challenge of limited interdisciplinary collaboration in time-series analysis due to data and methodological fragmentation across scientific domains.
  • To create a dynamic, self-organizing data repository that automatically identifies and highlights meaningful similarities between diverse time-series data, regardless of origin or metadata.
  • To lower barriers to data sharing by providing users with immediate, actionable insights into how their data relates to other empirical and synthetic systems.
  • To support the objective evaluation and development of time-series analysis algorithms by offering a large, diverse, and systematically organized dataset.
  • To serve as a template for future self-organizing data libraries across other data types, such as networks, images, and multivariate datasets.

Proposed method

  • CompEngine uses a canonical set of 185 time-series features to characterize each data object, capturing statistical, dynamical, and spectral properties.
  • These features are embedded into a low-dimensional space using t-SNE dimensionality reduction, enabling visualization and similarity computation.
  • Nearest-neighbor matching is performed in the feature space to identify time series with similar empirical dynamics, regardless of source or metadata.
  • The system supports real-time updates: as new data is uploaded, the library reorganizes and alerts users to new matches based on similarity.
  • Users can upload data with metadata (e.g., system type, recording method), and the platform automatically computes and displays interdisciplinary connections.
  • The platform is built on a scalable backend with database and web infrastructure, enabling persistent storage and interactive exploration of over 24,000 time series.

Experimental results

Research questions

  • RQ1How can time-series data from diverse scientific domains be meaningfully compared and connected despite differences in origin, sampling rate, and duration?
  • RQ2What types of interdisciplinary connections emerge when time-series data from biological, physical, financial, and synthetic sources are embedded in a common feature space?
  • RQ3Can a self-organizing data library improve data sharing and collaboration by revealing hidden similarities between seemingly unrelated systems?
  • RQ4To what extent can such a system support the objective evaluation and development of time-series analysis algorithms across diverse data types?
  • RQ5Can a feature-based, dynamics-driven approach to data organization outperform traditional metadata-based repositories in fostering scientific discovery?

Key findings

  • CompEngine organizes over 24,000 time-series data objects from diverse domains—including ECG, birdsong, financial data, and model simulations—into a shared feature space using 185 canonical features.
  • The t-SNE visualization reveals that time series from similar categories (e.g., gait, tremor, ECG) cluster together, while distinct categories remain spatially separated, validating the method’s ability to capture meaningful structure.
  • The system automatically identifies surprising interdisciplinary matches, such as connections between real-world biological dynamics and synthetic model systems (e.g., SDEs, oscillatory maps), suggesting shared underlying mechanisms.
  • Users receive real-time alerts when new data with similar features is uploaded, enabling ongoing discovery and collaboration as the library grows.
  • The platform enables comprehensive, bias-free evaluation of time-series algorithms by providing access to a diverse, empirically grounded dataset spanning multiple scientific domains.
  • CompEngine demonstrates the feasibility of a self-organizing data library that incentivizes data sharing not by metadata, but by revealing intrinsic, data-driven connections across disciplines.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.