5 ms·
> My personal experience is that InnoDB performance drops off a cliff as tables grow and parts of tables stop fitting in RAM. TokuDB just keeps on going. Yes b
by fweespeech 9y ago
> My personal experience is that InnoDB performance drops off a cliff as tables grow and parts of tables stop fitting in RAM. TokuDB just keeps on going.
Yes but at least in my experience as things fit in RAM, InnoDB is still the correct choice.
TokuDB is great once you run out of RAM-related solutions but given the price point RAM is at these days, we are talking 100GB+ tables which probably shouldn't be one a single machine anyway.
- deleted 9y ago[deleted]
- willvarfar 9y agoHa! A few gb ram and a few tb disk is real cheap in the cloud. If your dataset fits in ram, and when it doesn't you have to buy more boxes with more ram rather than just buying more disk, then I think you are making a very uneconomical trade off.
- fweespeech 9y ago> Ha! A few gb ram and a few tb disk is real cheap in the cloud. > If your dataset fits in ram, and when it doesn't you have to buy more boxes with more ram rather than just buying more disk, then I think you are making a very uneconomical trade off. It is all in what lower latency is worth to the bottom line and how well you scale horizontally.
- willvarfar 9y agoMy link described performance too. A cluster of innodb boxes is going to be slower than a single tokudb box. (Src: have run each, and stuff like shard-query to parallelise queries and scale them out)
- fgonzag 9y agoWhy wouldn't 100GB tables be in a single machine? It's not that much data. I'm comfortably running a 3TB table on PostgreSQL, which is growing at a rate of about 75GB a month. The machine isn't even that beefy, 128 GB of RAM, dual mid range Xeons, and prosumer SSDs in RAID 10. I would rather spend money in PCI-E SSDs, 300+ GB of RAM, and high end CPUs if it meant I didn't have to implement sharding at an application level.
- gaius 9y agowe are talking 100GB+ tables which probably shouldn't be one a single machine anyway. You can comfortably handle multi-tera tables on single machines these days. Adding more machines isn't "scaling", its admitting "we can't scale in software so make it a hardware problem"