5 ms·
This will be obvious, but I work at Tokutek. > When you have so many writes that sharding isn't enough. TokuDB does indexed insertions very fast. [1] We've e
by leif 14y ago
This will be obvious, but I work at Tokutek.
> When you have so many writes that sharding isn't enough.
TokuDB does indexed insertions very fast. [1] We've even plugged ourselves in underneath MongoDB (just for fun) and we beat them too. [2]
> When you make changes so fast you need a liquid schema.
TokuDB supports lots of schema changes with zero downtime. [3]
> When you want to make your boss learn map-reduce so he can query the data.
Can't help you there, but I can make you not need to torture your boss that way.
> When the application can care of integration and not the database.
What if it didn't need to?
We also get fantastic compression [4], retain full transactional semantics, and lots of other fun stuff.
Email us if you're curious!
[1]: http://www.tokutek.com/resources/benchmark-results/benchmarks-vs-innodb-hdds/#iiBench http://www.tokutek.com/resources/benchmark-results/benchmark...
[2]: http://www.tokutek.com/2012/08/10x-insertion-performance-increase-for-mongodb-with-fractal-tree-indexes/ http://www.tokutek.com/2012/08/10x-insertion-performance-inc...
[3]: http://www.tokutek.com/resources/benchmark-results/benchmarks-vs-innodb-hdds/#HotSchema http://www.tokutek.com/resources/benchmark-results/benchmark...
[4]: http://www.tokutek.com/resources/benchmark-results/benchmarks-vs-innodb-hdds/#Compression http://www.tokutek.com/resources/benchmark-results/benchmark...
- DanWaterworth 14y agoIn my experience, it's not just throughput that's important, but also 99th percentile latency. If I understand fractal trees correctly, you sometimes need to rewrite all of your elements on disk. How do you do this without causing lag?
- leif 14y agoThat's happily not how fractal trees work, at all. We have a few talks online describing how they work. Zardosht has one here http://vimeo.com/m/26471692 http://vimeo.com/m/26471692 . I thought Bradley had a more detailed one at MIT but I can't find it right now. (EDIT: found it! http://video.mit.edu/watch/lecture-19-how-tokudb-fractal-tree-indexes-work-1361/ http://video.mit.edu/watch/lecture-19-how-tokudb-fractal-tre...) Basically, I think you're thinking of the COLA. What we implement does have a literal tree structure, with nodes and children and the whole thing, so at any point you're just writing out new copies of individual nodes, which are on the order of a megabyte. At no point do we have to rewrite a large portion of the tree, so there aren't any latency issues.
- DanWaterworth 14y agoThank you very much for the links. Is my understanding correct? Essentially, a fractal tree is a B-Tree (or perhaps a B+Tree?) with buffers on each branch (per child). Operations get added to these buffers and when one becomes full, the operations get passed to the corresponding child node. Operations are applied when they reach the node that is responsible for the data concerned.
- leif 14y agoThat's about it! It's like a B+Tree in that the data resides in the leaves (well, and in the buffers), except that the fanout is a lot lower because we save room in the internal nodes for the buffers.
- DanWaterworth 14y agoI'm a little confused. I thought fractal trees worked cache obliviously or am I mistaken?
- leif 14y agoIn theory, it is cache oblivious, and the CO-DAM model informs our decisions about the implementation, but no, the implementation itself isn't actually cache oblivious. Shh, don't tell on us. A "fractal tree" is defined by our marketing team as "whatever it is we actually implement." If you want to talk cache-oblivious data structures, we can talk about things like the COLA and cache-oblivious streaming B-trees which have rigorous definitions in the literature. At some point, if you want to achieve a certain level of detail, you have to pick one or the other in order to continue the conversation.
- alex137 14y agoHi. I thought the block size is much bigger in fractal trees (like 4MiB instead of 4KiB) than in B-trees hence the fanout would be about the same? I'm trying to experiment with these ideas on my side, but can't quite grok how large the buffers must be at each level of the tree. Let's take a concrete example like: 2^40 (1T) records with an 8-byte key and an 8-byte value. In a traditional B-tree with a 8KiB block size and assuming links take 8 bytes too and assume for a while that blocks are completely full. So, that's 1G leaf blocks, a fanout of about 1K and hence & full 4-level trees: 1 root block, 1K level-2 blocks, 1M level-3 blocks, 1G leaf blocks. In this setting, I understand that a fractal tree will attach a buffer to each internal node. How large will they be at level 1 (root), level 2, and level 3 ?