4 ms·
While I agree that this configuration needs more memory, mainframes are primarily used for very high throughput jobs. "Lots of memory to sit idle" isn't really
by evol262 5y ago
While I agree that this configuration needs more memory, mainframes are primarily used for very high throughput jobs. "Lots of memory to sit idle" isn't really a priority.
That aside, samepage sharing has been a thing in virtualization for more than a decade. If you have 800 VMs running RHEL8 for POWER, a significant amount of the memory load is kept in identical pages, which lowers the burden of oversubscription. The recommended workload is unlikely to be 1000 unicorn/pet VMs, and more likely to be a large number of VMs running the same base OS, and some of the same workloads.
Synchronous interrupt locking is also a solved problem.
- jfindley 5y agoIf by "samepage merging" you mean KSM[0], it's hugely expensive in CPU cycles (not to mention severe impacts on membw, l3 cache, etc) and outside of low-end VPS hosting has relatively few real-world applications. I'm not sure why it would be a useful thing to mention here. 0: https://www.kernel.org/doc/html/latest/admin-guide/mm/ksm.html https://www.kernel.org/doc/html/latest/admin-guide/mm/ksm.ht...
- evol262 5y agoYes, I mean KSM (and its equivalents in other hypervisors). The CPU cycles are a non-issue for the configuration on this mainframe, and it's fair to assume that, since IBM developed the silicon and IBM engineers wrote the hypervisor layer, that most of the detrimental effects of "we need this to work across multiple CPU vendors and multiple generations of the architecture" are covered. Even x86 CPUs get improvements to cache coherency in memory dedupe scenarios with generational improvement, and intelligent NUMA topology layout helps a lot. It's also worth noting that it's a mainframe, so essentially every configuration has a large number of other processors with a couple of instructions disabled so they don't "count" as CPUs for licensing, but they still perform work. I would not be surprised in the least if there were 1+ co-processors dedicated to offloading operations exactly like "LPAR || z/VM memory dedupe"
- jfindley 5y agoThat's a good point - I was speaking from experience with KSM on x86, but having a dedicated coprocessor on system Z for this application would make a lot of sense and fix many of the problems I've seen on x86.
- pgtan 5y ago[edit] sorry, I've mistaken the concepts. The name is active memory deduplication, which sits on top on memory sharing. http://www.redbooks.ibm.com/redpapers/pdfs/redp4827.pdf http://www.redbooks.ibm.com/redpapers/pdfs/redp4827.pdf [old] It is called active memory sharing in the POWER world https://www.ibm.com/docs/en/power9/9080-M9S?topic=sharing-managing-powervm-active-memory https://www.ibm.com/docs/en/power9/9080-M9S?topic=sharing-ma... http://www.redbooks.ibm.com/redpapers/pdfs/redp4470.pdf http://www.redbooks.ibm.com/redpapers/pdfs/redp4470.pdf
- eptcyka 5y agoThis isn't something to be done at the guest VM level but at the hypervisor level. In an optimal scenario, it can be cheaper to clone a VM with CoW semantics than it would be to deduplicate memory dynamically. But I guess this would defeat kASLR between all of the clones.
- tw04 5y agoThe entire scenario outlined is for a production deployment of Websphere backed by Oracle. IBM specifically calls out 1:1 memory for production workloads (which oracle very much is) in their performance documentation: Slide 17: http://www.vm.ibm.com/library/presentations/syslimit.pdf http://www.vm.ibm.com/library/presentations/syslimit.pdf >Synchronous interrupt locking is also a solved problem. No, it really isn't. I have not seen a production implementation of virtualization with > 4:1 CPU oversubscription that doesn't eventually or immediately have significant performance issues with database workloads.
- zozbot234 5y agoIsn't WebSphere a web application server? That really doesn't look like the sort of extreme vertical scaling and integration that mainframes are supposed to perform well at. We would expect hyperscaled commodity hardware to perform best here.
- evol262 5y ago> The entire scenario outlined is for a production deployment of Websphere backed by Oracle. IBM specifically calls out 1:1 memory for production workloads (which oracle very much is) in their performance documentation: This discussion is about the configuration of the systems, not this specific scenario. While this may have been the scenario outlined for you, it wasn't apparent from your comment. Production workloads on busy servers have very different requirements than consolidation, VDI, resiliency/redundancy, hardware abstraction, etc. Every workload needs a different evaluation. All of these VMs running production Ora/OAM/OEM/Websphere workloads? No, don't oversubscribe. Some of them running similar workloads? More oversubscribe is ok. Few of them? Lots of oversubscribe. Similar for interrupt locking. The "old" synchronous interrupt lock I was speaking about was "I have a bunch of VMs with CPU oversubscribe, and even if they're doing nothing, 30% or more of your CPU time goes to interrupt scheduling so every vCPU for a given VM can schedule simultaneously". This is solved. "I'm running a CPU-intensive workload on a massively oversubscribed server" is not, and we really shouldn't expect it to be. I am having a generalized discussion about virtualization oversubscribe. You are having a specific discussion about CPU-heavy DB workloads. Apples cannot be compared to oranges.