Abstract
Time-series databases (denoted as TSDB), which are designed for handling rapidly growing time-series data, usually apply compression techniques to reduce storage overhead. However, existing compressors are limited in key metrics for TSDB compression, such as compression ratio or decompression speed, primarily due to a mismatch between their design and the features specific to TSDB. To this end, we propose a lossy compressor Machete. It achieves a much higher compression ratio and fast decompression speed, while promising a user-specific and point-wise error bound to preserve the analytical value of the data. First, Machete proposes a pattern-based predictor and an efficient hybrid encoder to monitor data trends, which successfully achieve higher compression rates through better understanding of the data. Second, Machete proposes a SIMD-based decompression acceleration technique, which exploits the repeatedly intermediate calculations in decompression and shares them in decompression iterations through parallelism. Our evaluation on four real-world datasets shows that Machete outperforms state-of-the-art compressors by 69%–114% on compression ratio and achieves the fastest decompression speed on two datasets. When applied to a well-known time series database InfluxDB, Machete saves disk usage 40%–72% and improves the query performance of the InfluxDB database by saving I/O.
| Original language | English |
|---|---|
| Article number | 131 |
| Journal | ACM Transactions on Architecture and Code Optimization |
| Volume | 22 |
| Issue number | 4 |
| DOIs | |
| State | Published - 13 Dec 2025 |
| Externally published | Yes |
Keywords
- IoT database
- lossy compression
- time-series data
Fingerprint
Dive into the research topics of 'The Design of an Efficient Lossy Compressor for Time Series Databases'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver