Towards optimal symbolization for time series comparisons

Smith, Gavin, Goulding, James and Barrack, Duncan (2013) Towards optimal symbolization for time series comparisons. In: IEEE 13th International Conference on Data Mining Workshops (ICDMW 2013), 7-10 Dec 2013, Dallas, Texas, USA.

Full text not available from this repository.

Abstract

The abundance and value of mining large time series data sets has long been acknowledged. Ubiquitous in fields ranging from astronomy, biology and web science the size and number of these datasets continues to increase, a situation exacerbated by the exponential growth of our digital footprints. The prevalence and potential utility of this data has led to a vast number of time-series data mining techniques, many of which require symbolization of the raw time series as a pre-processing step for which a number of well used, pre-existing approaches from the literature are typically employed. In this work we note that these standard approaches are sub-optimal in (at least) the broad application area of time series comparison leading to unnecessary data corruption and potential performance loss before any real data mining takes place. Addressing this we present a novel quantizer based upon optimization of comparison fidelity and a computationally tractable algorithm for its implementation on big datasets. We demonstrate empirically that our new approach provides a statistically significant reduction in the amount of error introduced by the symbolization process compared to current state-of-the-art. The approach therefore provides a more accurate input for the vast number of data mining techniques in the literature, providing the potential of increased real world performance across a wide range of existing data mining algorithms and applications.

Item Type: Conference or Workshop Item (Paper)
RIS ID: https://nottingham-repository.worktribe.com/output/720523
Additional Information: © 2013 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. Published in the Proceedings of the IEEE International Conference on Data Mining Workshops (ICDMW 2014)
Keywords: Time series analysis; Quantization (signal); Equations; Mathematical model; Data mining; Approximation methods; Simulated annealing
Schools/Departments: University of Nottingham, UK > Faculty of Social Sciences > Nottingham University Business School
Identification Number: 10.1109/ICDMW.2013.59
Depositing User: Eprints, Support
Date Deposited: 04 Jun 2018 08:05
Last Modified: 04 May 2020 16:40
URI: https://eprints.nottingham.ac.uk/id/eprint/52220

Actions (Archive Staff Only)

Edit View Edit View