5 ms·
Timeseries data storage in MongoDB
- iskander 15y agoAre there advantages over storing data in HDF? I've been working with a few hundred gigabytes of financial data this summer and I'm finding that python's data-oriented libraries (h5py, numpy, scipy, matplotlib, scikits.learn) cover my needs.
- seigenblues 15y agodepends on your usage. The rest of the toolchain is all python, numpy, scipy, matplotlib especially. This might be poorly titled: it's not just about the storage as it is about the aggregation of disparate sensor data into coherent, continuous data streams.
- iskander 15y agoI'm totally ignorant of mongodb: what does it do for you (in the way of data aggregation) that's not easy in numpy?
- hogu 15y agoif your data can fit into arrays, then there's no advantage in terms of the types of aggregations. however mongo allows you to store complex structures, think nested dictionaries/lists, and query on those nested structures, even allowing you to reach inside of nested structures to do the querying. I guess you could do nested structured arrays in numpy, I've never done that before.
- iskander 15y agoI use h5py's datasets (which are organized hierarchically and stored in compressed chunks) to do basic filtering and then load a fraction of my data into memory as numpy arrays.
- jorsh 15y agoNothing turns off my brain to listening to anything you've got to say faster than using reddit meme faces in your presentation slides.
- yannis 15y agoInteresting presentation, would do better with some more details in a blog or pdf.
- seigenblues 15y agothanks, i'll try and write them up. any particular areas you'd like to see expanded?
- yannis 15y agoI am particularly interested in the software implementation. I am a Mechanical Engineer, involved with high rise construction. Looking to disrupt Building Management Systems.
- theatrus2 15y agoOvertaking a BMS has huge problems (I know, I am in that space). On one end you're talking about certified hardware for a variety of needs (hardware, especially certified to various ASHRAE/ANSI/etc specs is expensive for a small company, no two ways about it). Second is the "no one got fired for buying IBM" mentality - if you hook up with Siemens, JCI and the like, you have paid that huge maintenance contract to make the vendor fix their problem and you won't end up with an insolvent small vendor's VAV controllers which don't work.
- yannis 15y agoThink positive, that is why you and I are on HN and our colleagues live in 30 year old specs space.
- cstuder 15y agoBeing in a similar industry, I share the sentiments of the last couple of slides: Dataloggers are expensive and horrible pieces of hardware. Proprietary solutions with no regard for real-life scenarios (Limited connectivity, power failures, connection failures, weird and inflexible data formats...) I would love to have a look at their Arduino based solution.
- moe 15y agoAw, this was physically painful to skim. What you really want for time-series data is a column db such as cassandra (or vertica etc.), perhaps HBase, perhaps a RDBMS, or perhaps a plain old log-file. What you most definitely don't want is Microsoft Access or MongoDB. Thinking about it, MS Access might still work to a degree.
- seigenblues 15y agoglad you didn't like it. Vertica was not free, HBase was really heavyweight, & I knew i didn't want RDBMS. The sensor data itself is plain csvs ;) Cassandra still does look interesting, although it looked like it would take me much longer to get it going; perhaps it will get rewritten. (The mongo solution only took a week to get something usable.)
- rbranson 15y agoCurious to hear why you thought it would take longer to get going with Cassandra. A single-node Cassandra cluster can be up with a 1-line command. Was it the availability of clients or interface abstractions? I understand that tutorial & documentation is not as readily available, so that's understandable.
- seigenblues 15y agoIt wasn't the ease of install or simplest-case-deployment, that's for sure. I did install it and kick the tires. This was around the end of 2009, and i don't think there were many clients available (i see pycassa dates to 2010-04). again, ease of development is much more important to us than performance; it'll probably be a long, long time before our db engine is the choke point, at which point we'll have resources to use a heavier tool.
- rbranson 15y agoThis makes sense, the landscape was very different. Cassandra wasn't something I would have used back then either.
- ericHosick 15y agoFor a temporal or time based key value store (I think this is kinda what the presentation shows) I used a collection that was something like: Temporal Collection { _id: "X1", data_temporal : [ { time_start: SomeDate, time_stop: SomeDate, _id: "ID2" }, { time_start: SomeDate2, time_stop: SomeDate2, _id: "ID2" }] Data Collection { _id: "ID1", parent: X1, data: { field1: "some info", field2: 34 }, _id: "ID2", parent: X1, data: { field1: "Some info new", field2: 34 } } What is cool about this is that if you have access to the data like ID1, you can easily find out when it was added and how it changed. If you have access to the temporal ID, X1, then at any time you can see what the data looked like. If you need to relate data, the "foreign key" used is the data_temporal ID. In this way, it is possible to ask what your key value store data looked like at any time. But, this could be off from the article. This also works quite well in a relational database.
- luigi 15y agoI saw this presentation live at MongoDC and it was awesome.
- snowwindwaves 15y agohe needed to find this board for his datalogger http://www.amazon.com/Webcontrol-Universal-Temperature-Humidity-Controller/dp/B001H4JXLU/ref=wl_it_dp_o?ie=UTF8&coliid=I2CE1S2ZFOUCP6&colid=2RTLJLUX5CSOZ http://www.amazon.com/Webcontrol-Universal-Temperature-Humid...
- amalag 15y agoWhat about Hadoop? You can also use Hadoop as a backend for hadoop style filesystems. What about using Hive? Does it also require a fixed schema?