10 ms·
Time Series, the new shiny?
- deleted 10y ago[deleted]
- Confusion 10y agoPoses the question So what’s the big deal? People have been recording temporally oriented data since we could chisel on tablets. Never answers it, but instead explains how Riak handles large time series. Certainly interesting, but I would like an answer to this question, as I don't understand the big deal.
- deleted 10y ago[deleted]
- ttctciyf 10y agoImmediately following the part you quote (my emphasis) Well, as it turns out, thanks to the software-eating- the-world thing and the Internet of Things we happen to be amassing *vast quantities* of all sorts of data [...] The demand for systems that are capable of storing and retrieving temporal data on an *ever increasing scale* necessitates systems that are specifically designed for this purpose. Strongly implying the big deal is the need to scale?
- siculars 10y agoThanks, ya I thought I answered my question there ;)
- Confusion 10y agoThat's an obstacle that needs to be overcome because you want to get somewhere. A challenging obstacle that has spawned an entire industry, but nevertheless not the goal. It is implied we can get somewhere new and exciting with these vastly larger timeseries. My question is: where?
- jholman 10y agoI think the big deal is allegedly "we now have so much data that we have a use for distributed databases with high performance", which allows them to go on and say "and our big deal is that we've built a high-performance distributed database tuned especially for time series". I only noticed it because I was shocked that there was an actual comprehensible value proposition in a blogpost about a NoSQL product.
- Twisell 10y agoHowever being system admin of a PostgreSQL database managing a lot of timeseries I was just like "meh use timestamp and good index, what's the deal?". Also for performance sake just cluster the data using the most used index... But yeah I guess if you really need NoSQL this should be nice. However sharding based on time will probably be one of the easiest approach for horizontal PostgreSQL deployment.
- sitkack 10y agoOne of the most compelling features of Riak, both KV and TS, is that it is masterless and replicated. When nodes fail for whatever reason, the cluster handles it transparently.
- deleted 10y ago[deleted]
- adamneilson 10y agoI think Datomic which has been doing time series data for years now answers the "what’s the big deal" question for me: "Given a value of the database, one can obtain another value of the database as-of, or since, a point in time in the past, or both (creating a windowed view of activity). These database values can be queried ordinarily, without having to make special queries parameterized by time." http://www.datomic.com/rationale.html http://www.datomic.com/rationale.html So in other words you get to see exactly how the whole database looked at any given moment through history (certain parallels with the blockchain I suppose).
- zcam 10y agoUntil you reach 10B datoms and then you're off for a fun ride sharding with multiple databases. Also in an IoT context or more general timeseries usage you very likely don't care about tx, but very much do about write throughput and not having a spof, which Datomic isn't good at at all. Datomic is an ok choice in some contexts, but the one detailed here is not one of them, you're likely to reach Datomic limits quickly and be in a world of hurt when you do.
- assface 10y ago> I think Datomic which has been doing time series data for years now answers the "what’s the big deal" question for me: This is known as "time travel queries". Postgres supported this back in the original version from the 1980s but then they took it out in 1997 because it takes up too much disk space. It's trivial to do in an MVCC system. You just turn off garbage collection (vacuuming in PSQL parlance).
- fit2rule 10y agoI've always found it quite curious that computer human interfaces have always focused on the noun/verb proposition of describing data, and not the time/place. Time is the only true constant in the universe, and yet computers are set up to track and control it, seemingly, as a second thought. Imagine if instead of having files/folders to (teach,confuse) Grandma, we simply had a time-based system of references. If Time was a principle unit of information that a user was required to understand as an abstract concept, I feel that it would result in far better user interfaces. We can see this in the Music-making world, where Time is the most significant domain over which a Musician exerts control. A DAW-like interface for managing events seems to me to be quite intuitive - for so many other non-musical applications - that its almost extraordinary that someone hasn't built an email system, or accounting system, or a graphical-design system, of applications, oriented around this aspect. (Of course, they are out there - but it seems that Time management makes the dividing line between "professional" and "dilettante" users rather thick...)
- saint-loup 10y agoRegarding time-based interfaces, consider the work of Gelertner (et al.). https://archive.wired.com/wired/archive/5.02/fflifestreams_pr.html https://archive.wired.com/wired/archive/5.02/fflifestreams_p... http://www.wired.com/2013/02/the-end-of-the-web-computers-and-search-as-we-know-it/ http://www.wired.com/2013/02/the-end-of-the-web-computers-an...
- pbowyer 10y ago> Imagine if instead of having files/folders to (teach,confuse) Grandma, we simply had a time-based system of references. If Time was a principle unit of information that a user was required to understand as an abstract concept, I feel that it would result in far better user interfaces. How would it be easier when I, let alone my Grandma, can't remember if I did something last week or the week before last? It seems it puts a higher cognitive load on the user.
- x5n1 10y agoYeah I actually don't think the human brain accounts very well for time. Like despite a decade has passed I don't feel as if time has moved very much for me. Despite the fact that people age, it does not seem an inbuilt thing to recognize your age. It just seems to be something that happens to your body. Thinking a bit more about it I don't think the brain accounts for time in long term memory, it does a better job in short term memory. That's why a musician can use time and we can't use it very well for stuff we stored a week ago.
- flatwhiskers 10y agoIn terms of getting data into RiakTS, would streaming something through Kafka be an option for instance?
- mdigan 10y agoHi, Basho employee here. Yes, Kafka is an option. Here's an example about using Kafka with Riak TS and Spark Streaming: http://docs.basho.com/riak/ts/1.3.0/add-ons/spark-riak-connector/usage/streaming-example/ http://docs.basho.com/riak/ts/1.3.0/add-ons/spark-riak-conne...
- the_alchemist 10y ago> Riak uses the SHA hash as its distribution mechanism and divides the output range of the SHA hash evenly amongst participating nodes in the cluster. Wait, Riak uses SHA as distribution hash? Why use a cryptographic hash for distribution and not something like Murmur3, if you're talking about high-performant[0] ? [0] http://blog.reverberate.org/2012/01/state-of-hash-functions-2012.html http://blog.reverberate.org/2012/01/state-of-hash-functions-...
- zeckalpha 10y agoMore time is likely spent synchronizing data across the network than hashing. (And the variance for the network time is likely high enough to account for the hash time)
- jnbiche 10y agoConjecture: not sure if this is the reason why SHA is used, but a useful side effect is that it may make users of Riak less vulnerable to certain types of denial of service attacks. Not 100% sure since I know little about how Riak works, but a more predictable hashing algorithm could make it easy for attackers to overload a given bucket with data, and slow down the db to a crawl.
- baq 10y agoi'd bet it doesn't matter.
- eggy 10y agoI don't know Riak, other than its a distributed NoSQL key-value data store. Time series has always been prevalent in the fintec and quantitative finance, and other disciplines for decades. I read a book in the early 1990s on music as time series data, financial tickers, and so on. How is Riak different, or more suited to use than Kdb + q, J with JDB (free), Jd (a commercial J database like Kdb/q)[2], or the new Kerf lang/db being developed by Kevin Lawler[3]? Kevin also wrote kona, an opensource version of the "K programming language"[4]. Kdb is very fast at time series analysis on large datasets, and has many years of proven value in the financial industry. [1] https://kx.com/ https://kx.com/ [2] http://www.jsoftware.com/jdhelp/overview.html http://www.jsoftware.com/jdhelp/overview.html [3] https://github.com/kevinlawler/kerf https://github.com/kevinlawler/kerf [4] https://github.com/kevinlawler/kona https://github.com/kevinlawler/kona
- MasterScrat 10y ago> How is Riak different, or more suited to use than Kdb + q, J with JDB (free), Jd (a commercial J database like Kdb/q)[2], or the new Kerf lang/db being developed by Kevin Lawler[3]? I would really be interested to see an informed answer to this question!
- anthonybsd 10y agoKDB is not distributed and K (APL) is not particularly pleasant to work with. While it has a proven track record in fintech no one I know of is particularly fond of working with this technology. Not to mention the cost of K developers (200K+). Riak is simply offering a free alternative to these systems that is very palatable.
- gricardo99 10y agoI’m not sure what you mean by "not distributed". With kdb+ you have a lot of flexibility in how to setup the database. You can organize the data to be stored in a distributed fashion (across multiple devices, multiple servers), you can setup query load balancers to distribute work-loads, and you can replicate to multiple servers/devices. You don’t have to use K, you code in q, which most people find far easier to read/write. There’s a wealth of information to help with all this[1], and a very responsive user group. But yes, you do need kdb+ expertise to get full use out of the tool. And yes, feelings seem to run strong towards kdb+, in both directions, love/hate it. And correct again, it’s not free and the licensing cost is definitely a hurdle to wider adoption. [1] - http://code.kx.com http://code.kx.com
- bbrazil 10y agoAre there performance numbers available? We're on the look out for suitable remote storage for prometheus.io, and would want to know the hardware that'd be required to handle 1M samples/s and how many bytes a sample takes up. It doesn't support full float64 which we need, but we could workaround by putting it into a 64 bit unsigned number.
- svjethani 10y agoWe have engaged a 3rd party that will be doing testing that will get published. Currently we have done testing around specific customer use cases.
- gnufied 10y agoLooks really nice. although I am bit sad to see that - it requires structured schema. I have been lookout for a metric collection system (like influxdb) and this would fit very well - except the schema part.
- siculars 10y agoYou can kinda fudge it by making one of the columns a varchar. It will store whatever you put in it like stringified json but not compute over it (arithmetic, filters, aggs). (author)
- rch 10y agoLove the simple install for development on a Mac. Thanks for that.
- siculars 10y agoI can't speak for the engineers but a lot of folks at basho are constantly building and tearing down riak single instance or multi instance clusters on redhat/ubuntu virtual machines. I do that and I also have different versions of riak sitting in their own folders on my mac hd. mac osx support +1.
- rdtsc 10y agoI see SQL support, that is interesting. Isn't Riak the premier NoSQL database. I guess it is a NoNoSQL db now ;-) The implementation of SQL part is so neat. Great work whoever did that. It uses yecc and leex that comes with Erlang and rebar even knows how to compile those. Very cool! https://github.com/basho/riak_ql https://github.com/basho/riak_ql
- gordonguthrie 10y ago[takes a bow]
- rdtsc 10y agoAwesome work!
- throwaway_xx9 10y agoAll non-trivial NoSQL databases support (or will support) SQL - otherwise programmers can insert data, but end-users can't consume it. No reporting, no revenue. Cassandra has actually deprecated their original data access API in favor of CQL (Cassandra SQL.)
- pg_is_a_butt 10y agoworst website design ever. couldn't read article.
- jsonninja 10y agoFor the TS experts out there, any real world experience with Influx? (https://influxdata.com/ https://influxdata.com/)
- thom_nic 10y agoI am not an expert but from using InfluxDB I think influx supports more features in terms of aggregation/rollups/ gap filling/ retention/ etc. But I suspect Riak TS would win on certain scalability use cases because it's based on the Dynamo architecture.
- thinkdevcode 10y agoFantastic db, but maybe not quite ready for production usage. We are using influx for a small portion of our ingestion engine as well as for storing server metrics. The updates/improvements have been pretty astounding over the last year, but also hard to keep up. I had to fork the nodejs library just to update it from 0.9 to 0.12 [0] because there were a LOT of breaking changes. Pre-0.9 there were many issues we ran into when it came to disk space & performance but they have all been resolved as of 0.9. Another thing to note (after speaking with them on a few occasions) is that they are only providing cluster support to their enterprise offering which wont be available till this summer. They do offer Relay which is their high availability tool for the open source version.[1] You really won't need clustering unless you're doing an insane amount of writes: single server performance is insane right now. We average 10k writes/sec with bursts up to 5x that and it doesn't break a sweat (on a cheap 2 cpu/7gb ram instance, with ssd block storage). [0] https://github.com/thinkdevcode/node-influx https://github.com/thinkdevcode/node-influx [1] https://docs.influxdata.com/influxdb/v0.12/high_availability/relay/ https://docs.influxdata.com/influxdb/v0.12/high_availability...
- pauldix 10y agoInfluxDB CTO here, thanks for the note! We've definitely had significant improvements even over the last two months. And the breaking API changes will be a thing of the past very soon: https://github.com/influxdata/influxdb/blob/master/CHANGELOG.md#v100-unreleased https://github.com/influxdata/influxdb/blob/master/CHANGELOG... 0.13 drops Thursday and 1.0 is the next release :)
- anandjoseph 10y agoPaysa[1] has a time series of companies ranked by talent density[2]. This shows how the quality of talent has changed in companies over time. [1] https://www.paysa.com https://www.paysa.com [2] https://www.paysa.com/company-rank https://www.paysa.com/company-rank
- epaulson 10y agoAs someone who deals with sensor data, the tricky part is really not the write-rate, but rather dealing with messy data. There's a lot of parallelism in sensor network streams, and for many domains you never look at the sensors from one device against the sensors of another device, so you can put them in entirely different databases and it doesn't matter. (It's not true in every case, of course, but if you're doing time series/streaming, ask yourself if it's true for you before picking a system) The real pain is handling data that arrives out of order or otherwise very late, or handling data that never arrives at all, or handling data that's clearly wrong. Worse, you may have streams that are defined/calculated from other streams for some algebra on series, e.g. series C is series A plus series B - so handling new data on A means you need to recalculate/update the view for C. Oh, and you'd like this all to be mostly declarative so you have some way to migrate between systems if you need to switch for whatever reason. Apache Beam/Google Dataflow gets a lot of this stuff right: it's not quite as declarative as I'd like but it gets the windowing flexibility right and handles restatements at a data model level.
- siculars 10y ago>The real pain is handling data that arrives out of order or otherwise very late Riak TS uses leveldb under the hood. Leveldb is natively sorted. In riak ts that includes the bucket. so the sort order is basically bucket/%PK where PK is your composite PK as defined in your CREATE TABLE statement. See Local Key [0]. [0] http://docs.basho.com/riak/ts/1.3.0/using/planning/ http://docs.basho.com/riak/ts/1.3.0/using/planning/
- siculars 10y agoHello, I'm the author of the post. Thanks for all the interest! AMA!
- im_down_w_otp 10y agoMight be worth looking into dalmatiner.io (DalmatinerDB) as an alternative to this. It's also built on riak_core to manage cluster membership and the top-level framework for dealing with routing and rebalancing. Waited for a long time for Riak TS to come out. Tried KairosDB & Cyanite, but the operational overhead of Cassandra wasn't something I wanted to buy into for such a narrow use case (infrastructure metrics store), and then suddenly out of nowhere DalmatinerDB was released. The code is clean, the architecture is solid, and the ops story is simple. I don't have any affiliation of any kind with the Dataloop folks. I am however a happy end-user. We do currently use Riak KV due to its CRDT support though.
- Licenser 10y agoDalmatinerDB has been around for a while now. I started working on it during the EUC in 2014, and it was released as part of ProjectFiFo the same year I just really suck at marketing ;). But I'm kind of curious, are you using Dalmatiner directly? And if so what is your use case?
- im_down_w_otp 10y agoCurrent deployment is Graphite (Carbon/Whisper) proper due to a lack of time to change it out, but eval'd Dalmatiner and should be migrating to it in the next 6 months (before the year is out). Made it through testing like a champ. The use case isn't particularly interesting though. Just multi-site machine and application metrics aggregation. Needed something that could work with Graphite tooling, would be highly-available, and be comparatively trivial to operate/maintain.
- walrus01 10y agoTime series does not necessarily have to be about 'huge' data either, just a much greater level of historical precision. Example: ISP sells a circuit with 95th percentile billing to a customer. If you poll SNMP data from a router interface on 60 second intervals and store it in an RRA file, you will lose a great deal of precision over time (because RRAs are highly compressed over time). You'll have no ability to go back and pull a query like "We want to see traffic stats for the DDoS this customer took at 9am on February 26th of last year". with time series statistics you can then feed it into tools such as grafana for visualization. An implementation such as openTSDB to grab the traffic stats for a particular SNMP OID and store it will allow you to store all traffic data forever and retrieve it as needed later on. The amount of data written per 60 second interval is miniscule, a server with a few hundred GB of SSD storage will be sufficient to store all traffic stats for relevant interfaces on core/agg routers for a fairly large sized ISP for several years.