2 ms·
As someone who deals with sensor data, the tricky part is really not the write-rate, but rather dealing with messy data. There's a lot of parallelism in sensor
by epaulson 10y ago
As someone who deals with sensor data, the tricky part is really not the write-rate, but rather dealing with messy data. There's a lot of parallelism in sensor network streams, and for many domains you never look at the sensors from one device against the sensors of another device, so you can put them in entirely different databases and it doesn't matter. (It's not true in every case, of course, but if you're doing time series/streaming, ask yourself if it's true for you before picking a system)
The real pain is handling data that arrives out of order or otherwise very late, or handling data that never arrives at all, or handling data that's clearly wrong. Worse, you may have streams that are defined/calculated from other streams for some algebra on series, e.g. series C is series A plus series B - so handling new data on A means you need to recalculate/update the view for C.
Oh, and you'd like this all to be mostly declarative so you have some way to migrate between systems if you need to switch for whatever reason.
Apache Beam/Google Dataflow gets a lot of this stuff right: it's not quite as declarative as I'd like but it gets the windowing flexibility right and handles restatements at a data model level.
- siculars 10y ago>The real pain is handling data that arrives out of order or otherwise very late Riak TS uses leveldb under the hood. Leveldb is natively sorted. In riak ts that includes the bucket. so the sort order is basically bucket/%PK where PK is your composite PK as defined in your CREATE TABLE statement. See Local Key [0]. [0] http://docs.basho.com/riak/ts/1.3.0/using/planning/ http://docs.basho.com/riak/ts/1.3.0/using/planning/