9 ms·
Computing Performance on the Horizon
- jeffbee 5y agoA delightful set of slides and references that tickles many of my pet topics. In particular, one I’d love to hear more about is why so many deployments are still choosing 2-socket servers by default when managing them is such a pain in the neck and the performance when you do it badly is so poor. Live the life of the future, today: choose single sockets!
- soulbadguy 5y agoIn world where most work loads are containerized, and where each container can be pinned to numa region doesn't it really matter?
- wmf 5y agoDoes any container runtime/orchestrator perform this optimization yet? Why wait?
- syoc 5y agoKernel scheduling is NUMA aware and will localize workloads. Threads will mostly have their RAM on the sticks local to their node. The core the thread is delegated to is also more likely to be the core local to the disk or NIC being used for IO. This is at least my experience, though I am no expert.
- solarkennedy 5y agoTitus (Netflix's container orchestrator that I work on) does this via: https://github.com/Netflix-Skunkworks/titus-isolate https://github.com/Netflix-Skunkworks/titus-isolate
- jeffbee 5y agok8s, by default, is oblivious to NUMA topology. You have to enable unreleased features and configure them correctly, which is the unwanted complexity to which I referred earlier. Simply aligning your containers to NUMA domains does not solve the problem that your arriving network frames or your NVMe completion queues can still be on the wrong domain. Isn't it simpler to just have 1 socket and not need to care? The number of cores available on a single socket system is pretty high these days, and in general the 1S parts are cheaper and faster.
- eloff 5y agoYeah, it makes a lot of sense to go with single socket servers unless you can't scale horizontally (e.g. database server). Why deal with the complexity when you can just side step it.
- dragontamer 5y agoWhy would you switch from a 100GBps NUMA connection (800 gigabits per second) over NUMA fabric into a 10 Gbps Ethernet fabric? If you are scaling horizontally, NUMA is the superior fabric than Ethernet or Infiniband (100Gbps) Horizontal scaling seems to favor NUMA. 1000 chips over Ethernet is less efficient than 500 dual socket nodes over Ethernet. Anything you can do over Ethernet seems easier and cheaper over NUMA instead.
- wmf 5y agoThis is correct if your software is NUMA-optimized (or if auto-NUMA works well for you) but if it isn't you can end up with slowdowns.
- dragontamer 5y agoSurely that can be fixed with just a well placed numactl command to set node affinity and CPU affinity. The root article is discussing rewriting code to fit on FPGAs. If NUMA is too complex then... I dunno. The FPGA argument seems dead on arrival.
- syoc 5y agoRack space can be quite expensive. Sometimes you need a lot of computing power in one or two rack units. Would be interested in what the management pains are. I agree that 2 socket machines require more thought in a lot of scenarios, especially IO heavy workloads.
- zaroth 5y agoIn my very limited experience it seems like space is much less an issue than power density. You can fit far more kW/U than the datacenter can possibly cool. In the commodity space that I rent, I ran out of power before filling even half the rack. I’m sure higher power/cooling density is possible to obtain, but I would think you’re primarily paying for that versus square footage?
- zozbot234 5y agoWhat do you need that power density for? It's a rack, not a supercomputer. (I sure hope it's not "mining coins" or anything like that.)
- wmf 5y agoIt's not really about needing power density; high density can easily happen accidentally. 40 1S servers in a rack could be 20 kW and 40 2S servers could be 30+ kW.
- deleted 5y ago[deleted]
- jeffbee 5y agoThe OpenCompute "Delta Lake" machine mentioned in the article occupies only one third of 1RU and peaks at 400W. You will certainly be power/cooling limited, rather than volume limited, with that kind of density.
- toast0 5y agoAll things considered, managing fewer hosts is nicer than more hosts. In some hardware generations, dual socket has pretty good cost and complexity tradeoffs. And if you benefit from having a large dataset in memory on a single machine, dual socket often gets you twice the DIMM sockets and therefore twice the ram. Quad socket has been very expensive (and not great performance) for quite some time, so that's usually out. Single socket Epyc looks pretty impressive though; although I'm retired and probably won't get to work with those anytime soon.
- rektide 5y ago> In particular, one I’d love to hear more about is why so many deployments are still choosing 2-socket servers by default bit of a guess but connectivity has been coming at a stupidly high premium for too long. there are sweet sweet blade chassis with price optimized less than full power 1p designs but 10g is still kind of novel there. bigger form factors are starting to see 25gbit at not-astronomical prices & switches in some rare cases are reasonable too. power supplies, storage, networking... a computer has a lot of not-entirely-ancillary needs. having multiple chips sharing the perhipetals should make sense, should be cheap. it's not though. the SMP tax is huge huge huge. thing is we don't need smp. we just need multihost peripherals. we needs nic's that like the grouphug ocp board can support 4 separate modes via pcie srv-io, a nice that can present multiple different virtual functions that different hosts can use. NVME similarly could be multiport- was was. power supplies are shared in ocp designs, with big bus rails, some 48v. I'm notad yet but sure seems dead obvious to me the future of multi-socket is non-coherency. build a big board with a couple different isolated computers on it, but connected via shared nic or nics to the top-of-rack. we get close with the 3 per width ope compute systems but those each need to be self contained, and there's an obvious leap in efficiency to be had by merging those three separate computers onto a single motherboard, while sharing some network, maybe storage devices. also like throw in some gratis pcie ntb maybe for a medium speed (32gbps on pcie 4.0 x16) direct server to server interconnect. ideally add another ntb unit on most chips so we can make a little mediums speed nearly free ring, or other topology. choose single sockets but choose many of them, each sharing some common peripherals.
- hackermeows 5y agoNice talk , covers a lot of base . He predicts Unikernels are dead, Containers will keep growing and lighthweight vms will take over after that.
- rektide 5y agoSome random contemporary musings, that touch some of these topics: I really hope we have a rad eBPF based QUIC/HTTP3 front-end/reverse-proxy router in the next 5 years. QUIC is so exciting and I just want it to be both fast & a supremely flexible way for a connection from a client to talk to a host of backend services. We'll definitely see some classic userland based approaches emerge, but gee, really hungry for For context, I was at the park two days ago, thinking about replacing a Node timesync[1] over websockets thing with a NTP-over-WebTransport (QUIC) implementation. There werent any H3 front-ends (which I kind of need because I just have some random colo & VPS boxes), and even if there were I was worried about adding latency (which a BPF based solution would significantly reduce, while letting me re-use ports 80/443). Especially as we see more extreme-throughput/HBM memory systems arrive, it's just so neat that we have a multiplexed transport protocol. Figuring out how to use that connection (semi stateless "connection", because QUIC is awsome) to talk to an array of services is an ultra-interesting challenge, and BPF sure seems like the go-to tech for routing & managing packets in the world today. QUIC, with it's multiplexing, adds the complexity that it is now subpackets that we want to route. I hope we can find a way to keep a lot of that processing in the kernel. [1] https://www.npmjs.com/package/timesync https://www.npmjs.com/package/timesync
- martinpw 5y agoSlide 26 is interesting - arguing that cloud providers have an advantage for future CPU design since they can analyze so many real world customer workloads directly. In previous roles I have worked with CPU vendors who have been very keen on getting access to profiling data from our workloads for design optimization, and lamenting the fact that it was hard to get such data and they were often limited to synthetic benchmark workloads when tuning new designs. So this argument does sound like a valid one, and does imply AWS etc will have significant advantages in future designs.
- handrous 5y agoAn interesting application of a now-familiar pattern: get lots of users, spy on them at massive scale, use those data to dominate some other market in a way that, at most, a single digit count of companies in the world could conceivably compete with (because none but they have anything like the data that you do). See also: everything to do with "AI".
- thechao 5y agoHe sort of implies that “just better hardware” will peter out in the 2030s. I think he’s calling it at least 50 years too soon. Here’s why: (1) I think logic designers are still faffing about in term of optimizing their designs; and (2), I think there’s a lot of smart people thinking “incrementally” through what we’d consider paradigm shifts in HW implementation. That is, our fabs will just naturally segue into 3D, spintronics, etc. I think he even mentions 3D circuits? One thing a lot of people miss is that layout of the design is materially different in 3D vs 2D: in 2D layout is NP-hard (complete) without efficient polynomial approximations; in 3D layout is low-order polynomial. The reduction in layout complexity will allow us to design things that are unthinkable right now, due to layout constraints & wire congestion.
- dragontamer 5y agoFPGAs from Xilinx are very complicated. They are no longer homogeneous 4-LUTs or 6LUTs with dedicated multipliers here and there. Today's FPGAs are VLIW minicores capable of SIMD execution with custom routing and some LUTs thrown around. They've stepped towards GPU style architecture while retaining the custom logic portions. FPGAs remain so difficult to use, I find it unlikely that they'd be mainstream in any capacity. GPUs seem like the easier way to get access to HBM + heavy compute, but either way the HBM future is eminent. ------------ GPUs have big questions about ease of use and practicality as it is, even with widespread acceptance of their compute potential. FPGAs are much less known, it's hard for me to imagine a mainstream future of them. Since memory bounds remains the biggest issue and not compute performance, I bet that the easiest to use accelerator with mass production and cheap access to the highest speed HBM is going to be the winner. GPUs are the current frontrunner, but the Fujitsu ARM CPU has easy access to HBM and could be a wildcard. POWER10 will be using high performance GDDR6. Not quite HBM, but it signals that IBM is also concerned with the memory bandwidth problem in the near future. CPUs could very well switch to HBM in some scenarios. ------------ If I were to guess the future: I think that AMD and NVidia have proven that today's systems need high speed routers to practically scale AMD has their IO die on EPYC. NVidia has NVLink and NVSwitch. That seems to be how to get more dies / sockets without additional NUMA hops. More efficient networks of chips with explicit switching / routing topologies is the only way to scale. The exact form of this network is still a mystery, but that's my big bet for the future. HBM is probably the future for high performance. DDR5 for cheaper bulk RAM but HBM on high performance CPUs / GPUs / FPGAs is going to be key. --------- The insight into RAM bottlenecks is interesting but seems to be point in favor of SMT. If your core is 50% waiting on RAM, then SMT into another thread to perform work while waiting on RAM.
- volta83 5y ago> If your core is 50% waiting on RAM, then SMT into another thread to perform work while waiting on RAM. If your core is 50% waiting on RAM, then SMT into another thread, and that other thread will want some memory to work on, so it will also wait on RAM. On Top of it, this second thread now puts extra pressure on the memory subsystem, might cause cache evictions for the other thread, etc etc etc The moment that you include the memory subsystem into the SMT picture, SMT goes from a "no brainer; waiting on memory? do other work" to a "uhhh... i don't know if this makes things better or worse".
- ksec 5y ago>for storage including new uses for 3D Xpoint as a 3D NAND accelerator; 3D XPoint's future is not entirely certain. Intel with their new CEO has remained rather quiet on the subject. Micron are pulling the plug on it and sold the Fab to Texas Instrument. The problem is there isn't a clear path forward with the technology, it make some sense when NAND and DRAM price were high in 2016 - 2019. Once they dropped to a normal level with newer DDR5 and faster SLC NAND or ZNAND with lower latency than XPoint's cost benefits becomes unclear. I guess we will know once Intel's Optane P5800X [1] is out with review. It is quite a beast. >Multi-Socket is Doomed Are there really no use-case where 128 Core+ with NUMA offer some advantage? >Slower Rotational Seagate [2] is actually working on dual Actuator HDD, think of it as something like internal RAID 0. The rational being as HDD gets bigger the time to fill up those drive increases as well. >ARM on Cloud Marvell partly confirms all HyperScalers have intention to build their own ARM CPU. But Google just announced their Tau instances [3], effectively cutting their cost / pref by 50%. Where each vCPU is an entire physical CPU core rather than a x86 thread. Not much mention on GPGPU. [1] https://www.intel.com/content/www/us/en/products/docs/memory-storage/solid-state-drives/data-center-ssds/optane-ssd-p5800x-p5801x-brief.html https://www.intel.com/content/www/us/en/products/docs/memory... [2] https://www.anandtech.com/show/16544/seagates-roadmap-120-tb-hdds https://www.anandtech.com/show/16544/seagates-roadmap-120-tb... [3] https://cloud.google.com/blog/products/compute/google-cloud-introduces-tau-vms https://cloud.google.com/blog/products/compute/google-cloud-...
- infogulch 5y ago> Are there really no use-case where 128 Core+ with NUMA offer some advantage? Are there any use cases where 128+ core single socket wouldn't be preferred to a 128+ core multiple socket design that is burdened by NUMA? AMD has been showing us that integrating the interconnects into the CPU package directly and letting it handle all the issues is a better design.
- dragontamer 5y agoWhen a hypothetical 128-core single socket comes out, will there be no workload that prefers to use a 2x128-core dual socket instead? AMD CPUs remain largely dual-socket compatible. Today's 64-core EPYCs can be dual-socketed into 2x64-core beasts. It just seems silly to me that if you're building say 200 computers in 10x racks (20-computers per 10x 40U racks) that you'd prefer single socket over dual-socket. If you're scaling up and out so much, what exactly is the problem with dual socket? Its not costs: dual socket remains cost-effective on a per-core basis over single-socket. Dual-sockets cuts the number of computers you need to work with in half. Etc. etc.
- rbanffy 5y agoI have an enormous respect for Brandon Gregg, but this "one socket ought to be enough for anyone" is something I saw too many people get burned with. I mean, it should, but who knows what the next version of Slack will need...
- hinkley 5y agoStill feels to me like we should be going the other way - kick more and more things off of the motherboard and support them with discrete - potentially customized - processors of their own. Between io_uring and current or future facilities of eBPF, we have a lot of tools on deck for pipelining IO operations, and once you have a way to pipeline IO operations, the latency is not the only bottleneck. Then it’s a matter of how much bandwidth you can push between two processes, or processors.
- nine_k 5y agoOK, it must be a compute-heavy load that does little random memory access, but works mostly on compact in-cache structures and maybe does sustained sequential memory accesses. With that, it's not suitable to offload to the GPU. What could it be? Serious question.
- dragontamer 5y agoDatabase with a large portion of data in-memory. * Second socket increases the memory channels and RAM available: 16-channel dual-EPYC with 8TB of RAM will be faster than 4TB of RAM on single-EPYC 8-channel. * SQL optimizers automatically search for sequential scans, because sequential scans are faster. * While JOIN can be done in GPU space, GPUs have extremely low memory capacity (only 80GB on the latest A100 that costs $10,000+). CPU will be faster because you can keep a much larger dataset hot in RAM. Your 80GB of VRAM on a GPU means nothing if your dataset is in the multi-TB range. (8TB of CPU-RAM on the other hand, serves as a reasonable cache)
- rbanffy 5y agoMore sockets add memory controllers, but we can also think about moving HBM closer to the cores as a L4 cache or scratch memory that’s not expected to be synchronised with other cores/sockets.
- ineedasername 5y ago3d CPU stacking seems interesting where surface area is a limited resource, but otherwise it seems like it would significantly complicate cooling things efficiently. Or isy assumption wrong?
- wmf 5y agoYou're right; you don't want to stack hot silicon on top of other hot silicon.