4 ms·
> We plan to deliver improvements to [..] purging mechanisms During my time at Facebook, I maintained a bunch of kernel patches to improve jemalloc purging mec
by adsharma 7mo ago
> We plan to deliver improvements to [..] purging mechanisms
During my time at Facebook, I maintained a bunch of kernel patches to improve jemalloc purging mechanisms. It wasn't popular in the kernel or the security community, but it was more efficient on benchmarks for sure.
Many programs run multiple threads, allocate in one and free in the other. Jemalloc's primary mechanism used to be: madvise the page back to the kernel and then have it allocate it in another thread's pool.
One problem: this involves zero'ing memory, which has an impact on cache locality and over all app performance. It's completely unnecessary if the page is being recirculated within the same security domain.
The problem was getting everyone to agree on what that security domain is, even if the mechanism was opt-in.
https://marc.info/?l=linux-kernel&m=132691299630179&w=2 https://marc.info/?l=linux-kernel&m=132691299630179&w=2
- jcalvinowens 7mo agoI'm really surprised to see you still hocking this. We did extensive benchmarking of HHVM with and without your patches, and they were proven to make no statistically significant difference in high level metrics. So we dropped them out of the kernel, and they never went back in. I don't doubt for a second you can come up with specific counterexamples and microbenchnarks which show benefit. But you were unable to show an advantage at the system level when challenged on it, and that's what matters.
- adsharma 7mo agoYou probably weren't there when servers were running for many days at a time. By the time you joined and benchmarked these systems, the continuous rolling deployment had taken over. If you're restarting the server every few hours, of course the memory fragmentation isn't much of an issue. > But you were unable to show an advantage at the system level when challenged on it, and that's what matters. You mean 5 years after I stopped working on the kernel and the underlying system had changed? I don't recall ever talking to you on the matter.
- jcalvinowens 7mo ago> By the time you joined and benchmarked these systems, the continuous rolling deployment had taken over Nope, I started in 2014. > I don't recall ever talking to you on the matter. I recall. You refused to believe the benchmark results and made me repeat the test, then stopped replying after I did :)
- nullpoint420 7mo agoThis is why I love hacker news. I learn so much from these moments.
- integricho 7mo agoI came here for the article, stayed for the drama.
- danudey 7mo agoLike "never work at Meta unless you can out-toxic your coworkers".
- __turbobrew__ 7mo agoYea I knew meta was toxic, but publicly beefing over something over a decade ago is a whole other matter. I can’t even remember what I was working on 10 years ago, and even if I did I wouldn’t be bringing people down that much later.
- baby 7mo agoThe problem is a lot of very strong engineers are also very difficult to work with. I worked at Meta too and can tell you the other side of the coin is that people who were too toxic could get canned as well!
- __turbobrew__ 7mo ago
- vardump 7mo agoI wouldn't be surprised if both 'adsharma' and 'jcalvinowens' were right, just at different points in time, perhaps in a bit different context. Things change.
- google234123 7mo agoI like your clocks!
- asveikau 7mo agoMaybe I'm misreading, but considering it OK to leak memory contents across a process boundary because it's within a cgroup sounds wild.
- adsharma 7mo agoIt wasn't any cgroup. If you put two untrusting processes in a memory cgroup, there is a lot that can go wrong. If you don't like the idea of memory cgroups as a security domain, you could tighten it to be a process. But kernel developers have been opposed to tracking pages on a per address space basis for a long time. On the other hand memory cgroup tracking happens by construction.
- asveikau 7mo ago> across a process boundary > within a cgroup Note the complementary language usage here. You seem to have interpreted that as me writing that it didn't matter what cgroup they are in, which is an odd thing to claim that I implied. I meant within the same cgroup obviously. Yes, you can read memory out of another process through other means.. but you shouldn't map pages, be able to read them and see what happened in another process. That's the wild part. It strikes me as asking for problems. I was unaware of MAP_UNINITIALIZED, support for which was disabled by default and for good reason. Seems like it was since removed.
- adsharma 7mo agoI was clarifying that there are CPU cgroups, network cgroups etc and the proposal touched only memory cgroups. The people deploying it are free to restrict the cgroup to one process before requesting MAP_UNINITIALIZED if there is a concern around security. At that point the memory cgroup becomes a way to get around the page tracking restriction. But I get why aesthetically this idea sounds icky to a lot of people.
- genxy 7mo agoWhat metrics were improved by your patches?
- adsharma 7mo agoSome more historical context. It wasn't a random optimization idea that I thought about in the shower and implemented the next day. Previous work on company wide profiling, where my contribution was low level perf_events plumbing: https://research.google/pubs/google-wide-profiling-a-continuous-profiling-infrastructure-for-data-centers/ https://research.google/pubs/google-wide-profiling-a-continu... https://engineering.fb.com/2025/01/21/production-engineering/strobelight-a-profiling-service-built-on-open-source-technology/ https://engineering.fb.com/2025/01/21/production-engineering... The profiling clearly showed kernel functions doing memzero at the top of the profiles which motivated the change. The performance impact (A/B testing and measuring the throughput) also showed a benefit at the point the change was committed. This was when "facebook" was a ~1GB ELF binary. https://en.wikipedia.org/wiki/HipHop_for_PHP https://en.wikipedia.org/wiki/HipHop_for_PHP The change stopped being impactful sometime after 2013, when a JIT replaced the transpiler. I'm guessing likely before 2016 when continuous deployment came into play. But that was continuously deploying PHP code, not HHVM itself. By the time the patches were reevaluated I was working on a Graph Database, which sounded a lot more interesting than going back to my old job function and defending a patch that may or may not be relevant. I'm still working on one. Guilty as charged of carrying ideas in my head for 10+ years and acting on them later. Link in my profile.
- genxy 7mo agoThis kind of thing always struck me as something that the MMU and the memory controller could team up on. When you give back memory, you could not refresh it for some cycles. Or you could DMA the same page of zeros over all of it, so the CPU isn't involved in menial labor.
- adsharma 7mo agoThis is an old debate that goes back 25+ years. One of the differences in how Linux and FreeBSD handle the issue. Linux developers believe that involving the CPU warms the caches and is a good thing.