2 ms·
Database servers like influxdb or druid provide flexibility at the cost of performance. If you want 3 million inserts per second and 14 million reads per second
by gsubes 10y ago
Database servers like influxdb or druid provide flexibility at the cost of performance. If you want 3 million inserts per second and 14 million reads per second you have to roll your own solution (like the one in the link). Though if you know a database that performs like that, I would appreciate enlightenment. :)
Any query language, network access, non-binary storage format will slow the data system down, so we had to create our own embedded timeseries system to get this speed.
- chaotic-good 10y agoI'm working on my own time-series database server here - https://github.com/akumuli/Akumuli https://github.com/akumuli/Akumuli The most limiting factor is not the query language or network its a storage engine itself (and not the HDD/SSD write speed). I've discovered that write speed in akumuli is limited mostly by compression algorithm speed and parallelization. Solution above is bad because they're have only time-based index. It's impossible to read only one time-series without decompressing entire chunks that contains a lot of data. This solution is write optimized but it will introduce a lot of read amplification. It's not suitable for interactive applications. I've done the same mistake too but moved away from this design.
- gsubes 10y agoThe use case is financial strategy backtests where normally you iterate through the whole dataset and only keep a lookback window in memory to access older data (one backtest iteration per thread per cpu core). Only occasionally having to jump to a completely new data window on demand. The only downside with the jumps is that they will have to load worst case 9999 data points to reach the desired 10000 index of the file chunk. But with most strategies this is no problem and can be fixed by increasing the in-memory lookback window on demand. Having smaller files for smaller chunks would have decreased read performance because of having to switch files too often. So in that regard the read amplification is not a real problem. And yes, currently the inmemory AHistoricalCache and file based ATimeSeriesDB is only for date keys, but could be made generic if desired (simple pull request). Anyway a full db server definitely has more things to account for than this custom solution for this specific problem. While the point here is to show that a custom made solution can be a lot faster than general purpose timeseries databases. Or do you have different opinions here?