6 ms·
Writing a Time Series Database from Scratch
- iksaif 9y agoYou may also want to check https://github.com/criteo/biggraphite/wiki/BigGraphite-Announcement https://github.com/criteo/biggraphite/wiki/BigGraphite-Annou... which is also about how to write a TSDB from Scatch but with different goals.
- bogomipz 9y agoI had a question about the following statement from the post: >"Prometheus's storage layer has historically shown outstanding performance, where a single server is able to ingest up to one million samples per second as several million time series" How are there one million samples per second equating to several million time series? Is a single sample not equivalent to a single data point in a time series db for a particular metric in Prometheus?
- ah- 9y agoI think this means that there are e.g. 10 million different time series, that each get a new sample appended every 10 seconds.
- bongonewhere 9y agoIs like everyone creating a time series database from scratch?
- rodionos 9y agoThe fact that many companies (FB, Uber, Google, Netflix, SO) roll their own TSDBs for metrics collection suggests that there is a real need. Or maybe there is not. It could a way to make boring system monitoring jobs fun and fancy again.
- metaobject 9y agoPerhaps these companies have such varied requirements that none of the existing TSDB systems fully support? We built our own custom time series db, along with a suite of tools for accessing, slicing, plotting, etc bc (at the time, at least) there was no support for bitpacking data, and storing/calculating certain spatial data operations.
- weq 9y agoMy first job was in 2005 writing a TSDB abstraction over sql server used for BI with defense. we merged ~30Gb of data per night in about 8hrs on a 32core machine which was severly limited by a platter SAN. We would get about 1 disk failure /6months. So do u need an actual TSDB? Or are other companies doing what we did?
- Rapzid 9y agoI was under the impression that influxdb's storage engine was going to be a viable solution for Prometheus at some point. Now they are writing their own; not sure where that interest went.
- sagichmal 9y agoThat was never true AFAIK. I recall one of the core devs did a comparison of storage engines years ago and found Influx unsuitable; that work led to the current-gen storage engine.
- bbrazil 9y agoThat was the 1st InfluxDB storage engine, things have evolved in the intervening years. The latest InfluxDB design is actually quite similar to the latest Prometheus design, what's different is our approaches to reliability and clustering. It's presently looking like Influx might once again be an option for long term storage for Prometheus.
- Rapzid 9y agoThat was my impression after the influxdb storage engine rewrite; I'm guessing the similarities are not a coincidence :) I'm not sure how interested influxdb would be to the idea, given their shrewd moves towards monetization, but it would be nice if the storage engine could be developed as a component of influxdb and adopted into Prometheus(ala rocks, level, etc).
- pas 9y agoThat was about remote storage. And that was around the time when Influx itself used pluggable backends. Now influx has its own storage engine, and prometheus too.
- ah- 9y agoExciting times in database land! It certainly seems like the good systems are converging on very similar storage architectures. This design is so similar to how Kafka and Kudu work internally. As the raw storage seems pretty optimal now, I suspect next we'll see a comeback of indices for more precise queries to get another jump in performance.
- jnordwick 9y agoI still can't figure out why people can't even come close to KDB+. It is a real conundrum. I've been waiting patiently for something to show up, but the gap seems to keep getting bigger instead of smaller. Is it that people want to make the problem more complex that it needs to be? Is it that those who know most about these issues don't share their secrets so implemented from the outside often don't have a good understanding of how to do things properly? If you were to asked the guy behind Prometheus if he's looked at the commercial offerings and what he's learned from them, would even be able to speak about them intelligently? There seems to be a huge skills gap on these things that I can't put my finger on. I'd love to be able to use a real TSDB, even at only half the speed and usefulness. It would be great for these smaller firms that cant or wont pay the license fees for a commercial offering until they get larger.
- rodionos 9y agoIn-memory, K/Q, expensive, old, closed source?
- angersock 9y ago> It would be great for these smaller firms that cant or wont pay the license fees for a commercial offering until they get larger. Why should we give that work away for free, especially if it's not that hard to roll your own for most small-scale needs?
- derefr 9y agoSmall company doesn't mean small scale, if the data is purchased/licensed instead of primary data. There are one-man startups trying to do build their business on analysis of Facebook-sized data-sets.
- angersock 9y ago> There are one-man startups trying to do build their business on analysis of Facebook-sized data-sets. Then maybe they should learn how to write their own databases, or pay people for their expertise. :)
- nicolaslem 9y agoThe description of this new storage engine does not explain how it manages the durability of the data. When you compare with the extreme efforts traditional databases take to ensure that unplugging a server will never ever result in data loss[0], silencing this problem makes me wonder. Is it that at this ingest rate even trying to ensure durability is a vain effort? [0] https://www.sqlite.org/atomiccommit.html https://www.sqlite.org/atomiccommit.html
- tveita 9y agoThey mention using a write-ahead-log, which should be sufficient for durability if implemented correctly.
- bbrazil 9y agoDurability is not a requirement in that sense. Consider that a regular scrape has happened and that data has been accepted by the DB but not yet flushed to disk. Whether the database dies just before or just after the scrape produces the same result: The data for that scrape isn't present when the server restarts. There plenty of other ways a scrape might not succeed that we have no control over (e.g. other end is overloaded, network blip), so there's not much point obsessing over this particular failure mode. > Is it that at this ingest rate even trying to ensure durability is a vain effort? It's not in vain, but it'd be a bad engineering tradeoff in terms of throughput.