4 ms·
My main concern for now is to be able to store/write the data collected from a million IoT devices in a scalable fashion, so there is not really a requirement f
by mads 9y ago
My main concern for now is to be able to store/write the data collected from a million IoT devices in a scalable fashion, so there is not really a requirement for what to do with it other than, when the data is there it will be analyzed on an adhoc basis and then we will see what we can do with it.
- brianwawok 9y agoThat is going to be hard to achieve. To know how to store the data, you need to know how you will query it. If you know how you will query it, you can say devise a way to store it in Cassandra.. which will scale up to PB. If you just throw in a PB of data, I am not aware of ANY system that is going to let you drop ad-hoc queries at the data and get fast answers. You effectively need to load most / much of the data off disk to process it. If you only want to store it and not query it yet.. store it off on flatfile. When you decide how you will query it, load it from the flatfile and switch writes to your new system.
- mads 9y agoI recently built a system capable of handling this load, but is was based on MongoDB and thats not an option in this project because of hardware constraints. Cassandra may be an option. I will look into it.
- ehllo 9y agoMaybe this is a better option for you. http://www.scylladb.com/ http://www.scylladb.com/
- johann8384 9y agoIf you were using OpenTSDB: 1 million devices, reporting once per minute. We'll round that up to 17k submissions per second. Let's say each submission is, um, 50 datapoints. That's 850,000 datapoints per second. Not really that hard to store. With 12 tags, OpenTSDB stores that in about 100 bytes per datapoint. You'd be writing 80MBps of data across you cluster. Not a big deal there. Let's say you have 12 nodes with 8 disks at 2TB per disk. 192TB with 3x HDFS replication, so 64TB of usable space. That gives you 828,504 minutes of capacity. 19 months or so. Queries will be fast also, but you need to set it up well, and, it's easy to do it wrong (sorry about that, we're making it better). I recommend Splicer with 4 or 5 query instances per data node to parallelize the queries and take advantage of locality (hit the query instances on the same node as the HBase region). I missed where the Petabyte of data came in, but I just made up some numbers here that are accurate from my experience.