7 ms·
> In our own testing, we have found that microbenchmarks can show an exaggerated impact. This. The talks about 30% (or even 300%!) impact based on a graph with
by renchap 9y ago
> In our own testing, we have found that microbenchmarks can show an exaggerated impact.
This. The talks about 30% (or even 300%!) impact based on a graph without any details or methodology, often itself based on a microbenchmark or a tweet that got viral is really not helping.
- Analemma_ 9y agoSome people really are seeing those slowdowns though, it all depends on your workload. If you’re one of the people slowed down by 50%, i imagine the fact that you’re an outlier is cold comfort.
- discoursism 9y agoBesides that one forum post yesterday from someone whose m1.medium instance slowed down, what other cases do we know of? If it's just a single person who has observed a slowdown, I'd say this worked out quite well overall.
- renchap 9y agoIt also depends on the platform you are on, the CPU vendor/model, the vulnerability you fix, and the patch you used. As mentioned in the article, there are 3 different (identified) vulns, each with its own set of patches, and sometimes multiples different ways to patch. Each of those patch will have its own impact, and work is currently in progress to find new ways to patch that will have a lower impact on performance. Also, this means we will need to be used to patch, as new vulns of this category will definitely be found, and new changes to kernels/cpu/compilers will be designed and will need to be applied.
- nikanj 9y agoMy app slows 10%, on an OS that slows 5%, running in a virtualization host that slows 10%. The numbers add up.
- Florin_Andrei 9y agoWell, it's not a slowdown really. The high performance of the physical CPU was achieved through means that compromise security. The top of the speed range was bullshit, basically. Now that the top is being lopped off, you experience the true speed of that CPU while operating in a secure manner.
- fvdessen 9y agoDo they ? If your usage is 50% userspace, 50% kernel, then having each slice slow by 10% does not result in a 20% slowdown.
- gpm 9y agoThey do... because they are measured in a weird way. If the usage is split 50/50 kernel/userspace. And the userspace changes slow the app down 10%. Then the user space changes slowed the userspace code down 20%. Likewise if the kernel changes slow the app down 10%, they slowed the kernel code down by 20%. So the total impact of the changes is 20% in your example.
- fvdessen 9y agoThat is indeed a weird way to measure and report. Without the associated userspace/kernel usage ratio those numbers are meaningless.
- acdha 9y ago> a tweet that got viral is really not helping. I especially liked the AWS forum post being breathlessly circulated which is from someone claiming that their business critical workload no longer fits on a single m1.medium (1 vCPU, <4GB RAM), joined by someone else who was actually swapping and hitting the OOM killer. There is a real impact but this is going to be the lazy IT worker's new plausible excuse of the day for months…
- floatingatoll 9y agoFor those whose workers try to use such an excuse, ask them for historical graphs showing that they designed and operated the system to use 99% of provisionable RAM, and ask them what emergency response planning they included with that design in the case that a fluctuation in available RAM occurred due to unforeseen circumstances.
- PuffinBlue 9y agoThat particular AWS forum post doesn't support your argument. Especially when Matt the AWS rep specifically acknowledged in that very thread that the update AWS had applied had had an effect on performance. It matters not what the workload was running on before or accusations of laziness, the key takeaway from that post was that there was an impact on performance caused by the Meltdown/Spectre mitigations. The fact another commenter barged into the thread with an unrelated issue simply shows that AWS forum operates within the Law of Forums (which states all threads are to be cross pollinated with unrelated issues from at least one other user). EDIT: I should say that if this fellow is the only affected user from all the computer users of the world, then the mitigations applied can be called successful. That seems unlikely :-)
- acdha 9y agoI think you missed my point, which was simply that any service which you care about should be running on n > 1 servers and with at least enough capacity planning so that a small percentage change in workload doesn't cause user-visible failures. That's especially true in an environment like EC2 where hosts fail and noisy neighbors can cause service degradation. In this case, the thread made it clear that the real problem was that they didn't have a good deployment story — notice how the original poster was mentioning needing to move everything to an m3.medium manually? I would be quite surprised if they haven't had other problems in the past (e.g. what happens when a system update kicks off if that server is already running at 90+% utilization? or when they get more users and/or the existing users start doing slightly more) but hadn't wanted to deal with the hassle of migrating.
- ghaff 9y agoI posted this elsewhere, but there's performance data from Red Hat Performance Engineering in this blog post. (Numbers for Linux obviously.) https://access.redhat.com/node/3307751 https://access.redhat.com/node/3307751
- Game_Ender 9y agoFor me that post is behind a RedHat subscriber paywall, is there a free link somewhere?
- colek42 9y agoFYI: You can see it with a free develper account It talks about a significant slowdown, 8%-19% for OLTP Workloads (tpc), sysbench, pgbench, netperf (< 256 byte), and fio (random I/O to NvME). 3%-7% for Database analytics, Decision Support System (DSS), and Java VMs
- ghaff 9y agoIt was pulled back to make a few changes. It's public again.
- lossolo 9y ago> Subscriber exclusive content Can you share more?
- ForHackernews 9y agoIt used to not be paywalled: > Measureable: 8-12% - Highly cached random memory, with buffered I/O, OLTP database workloads, and benchmarks with high kernel-to-user space transitions are impacted between 8-12%. Examples include Oracle OLTP (tpm), MariaBD (sysbench), Postgres(pgbench), netperf (< 256 byte), fio (random IO to NvME). > Modest: 3-7% - Database analytics, Decision Support System (DSS), and Java VMs are impacted less than the “Measureable” category. These applications may have significant sequential disk or network traffic, but kernel/device drivers are able to aggregate requests to moderate level of kernel-to-user transitions. Examples include SPECjbb2005 w/ucode and SQLserver, and MongoDB.
- rdtsc 9y ago> This. Sometimes this, sometimes that, it depends really. > The talks about 30% (or even 300%!) impact based on a graph without any details or methodology, often itself based on a microbenchmark or a tweet that got viral is really not helping. Of course it's helping, everyone learned about CPU architecture, speculative execution, caches, and started trading AMD and INTC all of the sudden. But to be serious, doesn't "details and methodology" apply to any benchmarks we see on Twitter. Or do people automatically start buying more servers when they read that tweet. Anyone here did that? The point of that number is warn people to go and measure their own workload. Google measure and yay! not much of an impact. It's great news. Someone who does telephony and maybe is sending lots of tiny little UDP packets back and forth might be impacted. Should they start crying and running around, no. They should measure because it might affect them.
- mastax 9y agoI don't care about incompetent corporate IT departments listening to tweets. But "regular people" will read articles quoting the 30% figure and disable Windows update because they don't want their computer to get 30% slower. I saw this happen on /r/pcgaming
- Krabby127 9y agoWhat do you think the actual average impact will be?