6 ms·
Its worth reading this follow-up LKML post by Andres Freund (who works on Postgres): https://lore.kernel.org/lkml/yr3inlzesdb45n6i6lpbimwr7b25kqkn37qzlvvzgad5hf
by lfittl 6mo ago
Its worth reading this follow-up LKML post by Andres Freund (who works on Postgres): https://lore.kernel.org/lkml/yr3inlzesdb45n6i6lpbimwr7b25kqkn37qzlvvzgad5hfd7ut@xv4cihno76wu/ https://lore.kernel.org/lkml/yr3inlzesdb45n6i6lpbimwr7b25kqk...
- jeffbee 6mo agoFunny how "use hugepages" is right there on the table and 99% of users ignore it.
- bombcar 6mo agoI’m absolutely flabbergasted by the performance left on the table; even by myself - just yesterday I learned Gentoo’s emerge can use git and be a billion times faster.
- globular-toast 6mo agoThe time spent by emerge is utterly dwarfed by the time spent to build the packages, so who cares? Maybe it's different if installing a binary system but don't think most people are doing that.
- LtdJorge 6mo agoWhen using multiple overlays, emerge-webrsync is ungodly slower compared to git.
- bombcar 6mo agoIf you can emerge in 2.86s user you can do it right before you emerge world, meaning it's all "done in one interaction" (even if the actual emerge takes an hour - you don't have to look at it. Whereas if emerge is taking 5-10 minutes, you have to remember to come back to it, or script it.
- account42 6mo agoThat's really not universally true. Building can be parallelized on modern multi-code CPUs (minus configure), emerge cannot and portage is really really slow.
- GandalfHN 6mo ago[dead]
- GandalfHN 6mo ago[dead]
- TacticalCoder 6mo agoAIUI in that thread they're saying "0.51x" the perf on a 96-core arm64 machine and they're also saying they cannot reproduce it on a 96-core amd64 machine. So it's not going to affect everybody both running PostgreSQL and upgrading to the latest kernel. Conditions seems to be: arm64, shitloads of core, kernel 7.0, current version of PostgreSQL. That is not going to be 100% of the installed PostgreSQL DBs out there in the wild when 7.0 lands in a few weeks.
- master_crab 6mo agoFor production Postgres, i would assume it’s close to almost no effect? If someone is running postgres in a serious backend environment, i doubt they are using Ubuntu or even touching 7.x for months (or years). It’ll be some flavor of Debian or Red Hat still on 6.x (maybe even 5?). Those same users won’t touch 7.x until there has been months of testing by distros.
- crcastle 6mo agoUbuntu is used in many serious backend environments. Heroku runs tens of thousands (if not more) instances of Ubuntu on its fleet. Or at least it did through the teens and early 2020s. https://devcenter.heroku.com/articles/stack https://devcenter.heroku.com/articles/stack
- nine_k 6mo agoDo they upgrade to the new LTS the day it is released?
- crcastle 6mo agoNot historically.
- rvnx 6mo agoand they are right, this is because a lot of junior sysadmins believe that newer = better. But the reality: a) may get irreversible upgrades (e.g. new underlying database structure) b) permanent worse performance / regression (e.g. iOS 26) c) added instability d) new security issues (litellm) e) time wasted migrating / debugging f) may need rewrite of consumers / users of APIs / sys calls g) potential new IP or licensing issues etc. A couple of the few reasons to upgrade something is: a) new features provide genuine comfort or performance upgrade (or... some revert) b) there is an extremely critical security issue c) you do not care about stability because reverting is uneventful and production impact is nil (e.g. Claude Code) but 99% of the time, if ain't broke, don't fix it. https://en.wikipedia.org/wiki/2024_CrowdStrike-related_IT_outages https://en.wikipedia.org/wiki/2024_CrowdStrike-related_IT_ou...
- justinclift 6mo agoNote that it's just not a single post, and there's additional further information in following the full thread. :)
- adrian_b 6mo agoYes, and in the following messages the conclusion was that the regression is mitigated when using huge pages.
- jeltz 6mo agoWhich you always should use anyway if you can.
- justinclift 6mo agoHmmm, it's not always that clear cut. For example, Redis officially advised people to disable it due to a latency impact: https://redis.io/docs/latest/operate/oss_and_stack/management/optimization/latency/#latency-induced-by-transparent-huge-pages https://redis.io/docs/latest/operate/oss_and_stack/managemen... Pretty sure Redis even outputs a warning to the logs upon startup when it detects hugepages are enabled. Note that I'm not a Redis expert, I just remember this from when I ran it as a dependency for other software I was using.
- fabian2k 6mo agoThat's transparent huge pages, which are also not the setting recommended for PostgreSQL.
- gmokki 6mo agoJava can work with transparent hugepages (in addition to preallocated hugepages), but you just use +AlwaysPreTouch to map them in during the startup so that at runtime there won't be any delays or jitter. Redis should add a similar option
- 6mo ago
- aftbit 6mo ago>If this somehow does end up being a reproducible performance issue (I still suspect something more complicated is going on), I don't see how userspace could be expected to mitigate a substantial perf regression in 7.0 that can only be mitigated by a default-off non-trivial functionality also introduced in 7.0.
- cr125rider 6mo agoThey said the magic words to get Linus to start flipping tables. Never break userspace. Unusably slow is broken
- anal_reactor 6mo ago> Maybe we should, but requiring the use of a new low level facility that was introduced in the 7.0 kernel, to address a regression that exists only in 7.0+, seems not great. Completely right. This sounds like a communication failure. Maybe Linux maintainers should pick a few applications that have "priority support" and problems with these applications are also problems with Linux itself. Breaking Postgres is a serious regression. Reminds me of a situation where Fedora couldn't be updated if you had Wine installed and one side of the argument was "user applications are user problem" while the other was "it's Wine, like come on".
- falcor84 6mo agoI for one liked the old and simple WE DO NOT BREAK USERSPACE attitude. https://linuxreviews.org/WE_DO_NOT_BREAK_USERSPACE https://linuxreviews.org/WE_DO_NOT_BREAK_USERSPACE
- reisse 6mo agoNot sure it is true anymore. I've encountered few userspace breaks in io_uring, at least.
- gcr 6mo agoPerformance regressions are different from ABI incompatibilities. If the kernel refused to do any work that slowed down any userspace program, the pace would go a lot slower.
- shadowgovt 6mo agoOr be a lot uglier. See: Microsoft replacing its own API surfaces with binary-compatible representations to workaround companies like Adobe adding perf improvements like bypassing the kernel-provided kernel object constructors because it saved them a few cycles to just hard-code the objects they wanted and memcpy them into existence.
- cogman10 6mo ago
- fxtentacle 6mo ago.. which confirms all of my stereotypes. Looks like the AWS engineer who reported it used a m8g.24xlarge instance with 384 GB of RAM, but somehow didn't know or care to enable huge pages. And once enabling them, the performance regression disappears.
- bushbaba 6mo agoBecause such settings aren’t obvious to those not familiar with them. LLMs should make discoverability easier though
- perrygeo 6mo agoHonest question: what's the value of running the benchmark and reporting a performance regression if the author is not familiar with basic operation of the software? I'd argue that not understanding those settings disqualifies you from making statements about it.
- cogman10 6mo agoThe performance was reduced without a settings change. That is still a regression even if huge pages mitigates the problem. I'd be curious to know if there's still a regression with hugepages turned on in older kernels. If you are benchmarking something and the only changed variable between benchmarks is the kernel, that is useful information. Even if your environment isn't correctly setup.
- justinclift 6mo agoSome software clearly wants hugepages disabled, so it's not always the slam dunk people seem to be making it out to be. ie Redis: https://redis.io/docs/latest/operate/oss_and_stack/management/optimization/latency/#latency-induced-by-transparent-huge-pages https://redis.io/docs/latest/operate/oss_and_stack/managemen...
- perrygeo 6mo agoYet we're talking about postgres, specifically. The whole point is that benchmarks about postgres better know how to configure postgres or their conclusions be irrelevant at best. What does redis have to do with this discussion?