5 ms·
TimeSeriesDB
- efoto 10y agoThe link doesn't seem to have anything to do with timeseriesDB.
- sciurus 10y agoThe section named ATimeSeriesDB talks about an approach to build a persistence layer for time series data. I agree, not a great submission, though.
- gsubes 10y agoWell, true the title could have been better chosen, forgot there was a timeseriesDB in .NET Though the original title still was true about the approach in the link being more than 30 times faster than a pure LevelDB solution (which is also often used for this sort of storage).
- comboy 10y agoSince we have this title, any recommendations for open source time series database? Regarding pruning I'd prefer something more configurable than simple round robin, and it would be great to be able to store things like events (maybe logs), not only floats over time.
- dlbucci 10y agoWe've been using Druid (http://druid.io http://druid.io) at work for just that and I think we're generally pretty happy with it. It's column-oriented though, so while it's good for metrics and aggregations, you're never really retrieve the whole JSON objects that you put into it.
- wspeirs 10y agohttps://influxdata.com/ https://influxdata.com/ - purpose built DB for storing time series data.
- gsubes 10y agoDatabase servers like influxdb or druid provide flexibility at the cost of performance. If you want 3 million inserts per second and 14 million reads per second you have to roll your own solution (like the one in the link). Though if you know a database that performs like that, I would appreciate enlightenment. :) Any query language, network access, non-binary storage format will slow the data system down, so we had to create our own embedded timeseries system to get this speed.
- chaotic-good 10y agoI'm working on my own time-series database server here - https://github.com/akumuli/Akumuli https://github.com/akumuli/Akumuli The most limiting factor is not the query language or network its a storage engine itself (and not the HDD/SSD write speed). I've discovered that write speed in akumuli is limited mostly by compression algorithm speed and parallelization. Solution above is bad because they're have only time-based index. It's impossible to read only one time-series without decompressing entire chunks that contains a lot of data. This solution is write optimized but it will introduce a lot of read amplification. It's not suitable for interactive applications. I've done the same mistake too but moved away from this design.
- gsubes 10y agoThe use case is financial strategy backtests where normally you iterate through the whole dataset and only keep a lookback window in memory to access older data (one backtest iteration per thread per cpu core). Only occasionally having to jump to a completely new data window on demand. The only downside with the jumps is that they will have to load worst case 9999 data points to reach the desired 10000 index of the file chunk. But with most strategies this is no problem and can be fixed by increasing the in-memory lookback window on demand. Having smaller files for smaller chunks would have decreased read performance because of having to switch files too often. So in that regard the read amplification is not a real problem. And yes, currently the inmemory AHistoricalCache and file based ATimeSeriesDB is only for date keys, but could be made generic if desired (simple pull request). Anyway a full db server definitely has more things to account for than this custom solution for this specific problem. While the point here is to show that a custom made solution can be a lot faster than general purpose timeseries databases. Or do you have different opinions here?
- daenney 10y agoLogs are very much not a time series so storing them in a time series database is not something that I would recommend. Time series are essentially a sequence of numeric data points in chronological order. Though events are usually in chronological order, they're not numeric data points (though they can contain that too). As a consequence of being two different things, storing and analysing them efficiently need different solutions. I've seen people stuff everything in Elastic Search and it's certainly possible if you really want to.
- bbrazil 10y agoA time series database is anything that stores data with time attached, so log storage does count. The more pertinent difference is event logging vs. metrics. For the former the ELK stack is popular, and Prometheus.io is suitable for the latter.
- eis 10y agoPrometheus unfortunately has limited long term storage support and it's explicitly not one of their goals. Their goal is "operational monitoring", not analytics. I found that by trying to see if it could be used as a metrics DB but stumbled upon a few issues like configuring retention times per target. See for example https://github.com/prometheus/prometheus/issues/1381 https://github.com/prometheus/prometheus/issues/1381 Prometheus has lots of potential but it's not a metrics DB at this point or in the near future. I just wish they'd made that a bit clearer on the page. "We make design decisions that presume that Promtheus data is ephemeral, and can be lost/blown away with no impact." That viewpoint pretty much limits them to be an ops tool. I wish they'd reconsider this point.
- Shish2k 10y ago> Though events are usually in chronological order, they're not numeric data points (though they can contain that too) I think the thing is that the primary use-case for a lot of people combines the two - eg they want to log a whole key:value map where values can be numeric or string, and then dynamically generate time series. We can do this with SQL, like if I want a graph of traffic per domain name, "SELECT time_bucket, domain_name, sum(bytes) FROM access_log GROUP BY time_bucket, domain_name" - but all SQL servers are totally non-optimised for this use case. Surely there must be some software out there which can do this efficiently?
- rz2k 10y agoHow about MonetDB? I think it was one of the earlier column store dbs, and its seems to remain actively developed open source, when others are proprietary products now.
- suls 10y agoKDB+ .. Not opensource but you can use the 32bit version for free
- deleted 10y ago[deleted]