[Paper Review] The UEA multivariate time series classification archive, 2018
This paper introduces the first UEA multivariate time series classification archive (2018) with 30 datasets, standardized equal-length formatting, and train/test splits to enable rigorous MTSC evaluation.
In 2002, the UCR time series classification archive was first released with sixteen datasets. It gradually expanded, until 2015 when it increased in size from 45 datasets to 85 datasets. In October 2018 more datasets were added, bringing the total to 128. The new archive contains a wide range of problems, including variable length series, but it still only contains univariate time series classification problems. One of the motivations for introducing the archive was to encourage researchers to perform a more rigorous evaluation of newly proposed time series classification (TSC) algorithms. It has worked: most recent research into TSC uses all 85 datasets to evaluate algorithmic advances. Research into multivariate time series classification, where more than one series are associated with each class label, is in a position where univariate TSC research was a decade ago. Algorithms are evaluated using very few datasets and claims of improvement are not based on statistical comparisons. We aim to address this problem by forming the first iteration of the MTSC archive, to be hosted at the website www.timeseriesclassification.com. Like the univariate archive, this formulation was a collaborative effort between researchers at the University of East Anglia (UEA) and the University of California, Riverside (UCR). The 2018 vintage consists of 30 datasets with a wide range of cases, dimensions and series lengths. For this first iteration of the archive we format all data to be of equal length, include no series with missing data and provide train/test splits.
Motivation & Objective
- Provide a public, standardized benchmark for multivariate time series classification (MTSC).
- Expand MTSC evaluation beyond small, domain-specific sets to encourage rigorous comparisons.
- Format data to equal length with no missing values and supply train/test splits for all problems.
- Host the archive and accompanying tools at timeseriesclassification.com to facilitate reuse by researchers.
- Categorize datasets into domains (HAR, Motion, ECG, EEG/MEG, Audio, etc.) and document data sources.
Proposed method
- Assemble the first iteration of the MTSC archive with 30 datasets spanning diverse domains.
- Standardize all data to equal length, remove missing data, and provide explicit train/test splits.
- Deliver data in Weka multi-instance format with per-dimension representations and relational attributes.
- Provide downloadable code to split multivariate ARFF files for flexibility across experiments.
- Bundle the entire archive (zip ~2GB) and host it at timeseriesclassification.com for easy access.
Experimental results
Research questions
- RQ1How many MTSC datasets are included in the 2018 UEA archive and what domains do they cover?
- RQ2What data formatting and preprocessing steps are used to standardize MTSC problems for fair comparison?
- RQ3How are train/test splits defined and provided for each dataset?
- RQ4What tooling is provided to manipulate and reuse the MTSC datasets (e.g., splitting ARFF files)?
Key findings
- The 2018 vintage contains 30 multivariate time series classification datasets.
- All problems are reformatted to equal length with no missing data and include train/test splits.
- The archive is available as a single ~2GB zip file with per-problem directories and Weka multi-instance format.
- The data are organized into domains such as Human Activity Recognition, Motion, ECG, EEG/MEG, Audio Spectra, and Others.
- Code to split multivariate ARFF files is provided to facilitate reuse across studies.
- The archive is hosted at www.timeseriesclassification.com for public access.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.