5 ms·
Gorilla: A Fast, Scalable, In-Memory Time Series Database [pdf]
- pdarshan 11y agoFew folks from Fb started this company called Interana, and they seem to be doing the same thing.
- nwmcsween 11y agoWhy pointers, why not just do a mirror mmap if you have constant offsets and if time points change and querying based on time points need be constant maybe a table that holds an offset w/ the difference? Also why not atomics instead of spinning?
- scurvy 11y agoI also saw spinlock and immediately thought, why?
- simpsond 11y agoThe compression is very neat.
- scurvy 11y agoA few things here: 1) What/where exactly are they using GlusterFS for? Has Gluster fixed their scaling problems yet? Specifically the issue where new storage spaces/nodes were only available to new directories and files, but not existing directories? Granted, the last time I looked at this was 2009 or so, but it was a flaw due to their "no master node" topology. 2) FB has an entire team to manage Hadoop/HBase. This shows just how much of a beast that stack is. Anyone who has run Hadoop on "Internet time" knows what I'm talking about. It's great at running time insensitive, deferred compute jobs in an academic or scientific setting. It's really hard to keep it all 100% running in an on-demand setting. Aside, I couldn't imagine just working on 1 product in an operations setting as my full-time job. Boredom/fatigue must be a problem on that team. 3) I'd like to see more information on the networking side. What transport protocol? How large are the average updates in frame size? Etc etc. We've built something similar to Gorilla in-house, so I'm happy to see that we've come to some of the same conclusions.
- vidarh 11y agoFor GlusterFS, new storage space is immediately available to new directories and files, and can be quickly made available for new files in existing directories by fixing the layout. You can then run a rebalance in the background if you want to evenly distribute the files, but that's of course a slow operation in a large cluster.
- nwmcsween 11y agoGlusterfs will just always be slower due to the userspace -> kernelspace -> userspace switching fuse has to do. I don't understand why facebook hasn't jumped on ceph yet.
- linuxhansl 11y agoFor #2, show me any system that holds over 2PB of data on a large set of machines that does not need a team to be managed.
- scurvy 11y agoI meant dedicated team. Most places run shared ops teams. Every place I've seen that runs a "big Hadoop" deployment also has a dedicated "Hadoop team" to go with it. It's pretty easy to build a 2PB storage system on Ceph that the average group of sysadmins can run.
- saosebastiao 11y agoI really wish this included a comparison with KDB. It's not cheap to get a license, and they certainly wouldn't give a testing license in order to publish benchmarks against it, but in finance it is the standard for TSDBs. There hasn't ever been anything open source that has come close.
- beagle3 11y agoThey can download a 32-bit version of kdb+ for evaluation purposes, not sure if they could publish any benchmarks, and it would obviously not properly represent the speed (and capacity) of the 64-bit version. But I suspect that even the 32-bit kdb+ is going to be significantly faster than this gorilla.
- iskander 11y agoI briefly worked with K and Q while doing research on high-level numerical computing. I found their claimed efficiency to be severely exaggerated. K is a very naively implemented array-oriented language and I found it to be slower than Matlab or NumPy for many tasks.
- tlack 11y agoYou should consider writing up something with examples. Given how it's implemented (unboxed directly typed mmap'd arrays) it's hard to see how much slack there could be in it. I and many others are beginning to experiment with Q and Kdb and your learnings might spark valuable debate - and perhaps saved time. :)
- iskander 11y agoI wouldn't voluntarily touch K/Q ever again, but if someone else is interested in doing a blog post: try implementing any iterative machine learning algorithms. The separate compilation of primitive operators requires creating many array temporaries that on the one hand, Matlab's JIT can fuse and, on the other hand, NumPy provides a richer set of compiled functions to work with. K/Q lets you (tersely) express complex array computations using just its core operators, but all those array constructions add up to comparatively bad performance.
- thrusong 11y agoSo this isn't managing news feed data or anything like that, it's helping them aggregate server performance and error data for quick look up?
- rodionos 11y agoHas anyone attended a VLDB conference recently? How is it different from Strata, for example? P.S. Their choice of venues is nice.
- rodionos 11y ago> Further, many data sources only store integers into ODS If the underlying data type is 64 bit double, aren't they losing precision for integers greater than 2^53?