3 ms·
In essence, although we had moved 5% of the data from shard0 to the new third shard, the data files, in their fragmented state, still needed the same amount of
by donaldc 16y ago
In essence, although we had moved 5% of the data from shard0 to the new third shard, the data files, in their fragmented state, still needed the same amount of RAM. This can be explained by the fact that Foursquare check-in documents are small (around 300 bytes each), so many of them can fit on a 4KB page. Removing 5% of these just made
each page a little more sparse, rather than removing pages
altogether.
Interestingly, this is one of the reasons antirez gives as to why redis will not be using the built-in OS paging system, but instead will use one custom-written for redis' needs.
- IgorPartola 16y agoI couldn't help but think of that exact issue. I suppose once compacting the data online is built in, this particular issue won't come up again. At the same time, when a machine is overloaded, often times you have even bigger problems. For example, if you are out of memory, you may not be able to create another SSH process to get at the box.
- kunley 16y agoThe situation antirez had in mind can be remedied by occasionally using a tool like vmtouch to steer what's in the OS cache.
- IgorPartola 16y agoWell, the other big advantage being that the OS has no idea about the particular kinds of data you want to store, whereas Redis has quite a bit more data to work with. But yes, vmtouch looks useful for these types of situations (thanks).
- deleted 16y ago[deleted]
- Nate75Sanders 16y agoWe need an "Ask HN" for cool tools. Maybe there's already a SO question about it. I hadn't heard of vmtouch and probably tons of other things people here use.
- kunley 16y agoThe funny yet encouraging thing is that I'm just passing the wisdom: I've heard of vmtouch here on HN :)
- megablast 16y agoSurely if they moved more than 5% across, this would have freed up more memory, despite the fragmentation. Maybe they would be better identifying regular users, these would require on the fly compacting, be hosted on one or two machines, with a third smaller server for the non-frequent users.
- bmm6o 16y agoFreeing memory would require that there be an entirely empty page. Each page holds 13-14 objects (4k/300). The chances that 13 consecutive objects were in the 5% (assuming independent, random distribution) is 1 in 20^13, putting the expected number of empty pages well below 1. You have to migrate much more data before you can hope to see page-size holes.
- wheels 16y agoI don't think those are the same problem -- could you provide a link? The problem isn't paging, per se -- the paging system is doing exactly what it should be doing, paging in blocks off of the disk and into memory as they become hot, which for this use case is always. The problem is that you get fragmentation in your pages. If you allocate three records in a row that are 300 bytes, and then need to rewrite the first one to make it 400 bytes, or delete it altogether, you end up creating a hole there. The typical strategy for dealing with those holes it to maintain a list of free blocks which can be used and hope that the distribution of incoming allocations neatly fits into unused chunks. However, as they note, if you have a fully compacted 64 GB active data set and remove 5% of it you just end up with address space that looks like swiss cheese; it's not set up to elegantly shrink, but to recycle space as it grows. There are a couple of things that I find a bit odd here though: they should have seen this coming before migrating data; it's pretty obvious from the architecture. Second is that they mention the solution being auto-compacting, which wouldn't have actually helped them. Auto-compaction is in fact useful, but all that it does is, well, compact stuff. It means they would have hit the limits later, but once the threshold was crossed, they'd have the exact same problem. Auto-compaction is either an offline process that runs in the background or a side-effect of smarter allocation algorithms. Both of those things need time once you remove data from an instance to reclaim the holes in the address / memory / disk space ... which is exactly what they did manually. The only really sane ways to handle something like this are notifications at appropriate levels, or block-aware data removal -- e.g. "give me stuff from the end of the file". I don't know if mongo uses continuation records and stuff like that enough to know how difficult that would be for them. (Note: Directed Edge's graph database uses a similar IO scheme, so I'm doing some projecting of our architecture onto theirs, but I assume that the problems are very similar.)
- donaldc 16y agoI don't think those are the same problem -- could you provide a link? You are correct, the actual problem is paging out LRU keys as opposed to memory holes. The issue is related but not the same. From http://antirez.com/post/what-is-wrong-with-2006-programming.html http://antirez.com/post/what-is-wrong-with-2006-programming.... Multiply this for all the keys you have in memory and try visualizing it in your mind: These are a lot of small objects. What happens is simple to explain, every single page of 4k will have a mix of many different values. For a page to be swapped on disk by the OS it requires that all contained objects should belong to rarely used keys. In practical terms the OS will not be able to swap a single page at all even if just 10% of the dataset is used.