4 ms·
> When you cut out all that fat, you can fit a lot more muscle in its place. Then you can arrange everything on a rack scale for airflow, power, redundancy, etc
by throwawaylinux 3y ago
> When you cut out all that fat, you can fit a lot more muscle in its place. Then you can arrange everything on a rack scale for airflow, power, redundancy, etc.
This is not really true. Data center servers are highly optimized for density already. Like tens of thousands of hours on airflow CFD, tweaking cases and fans and internal layout, heatsink design, etc. chasing the 0.1%s. There are many significant tradeoffs to be made, but there is not a large factor improvement just sitting on the shelf due to a bunch of unused legacy leftovers.
- goombanast 3y agoPeople have this idea that legacy code or hardware is some huge burden but due to continuous progress in scale the relative size of legacy code or hardware shrinks over time. There might be some exception (e.g. physical connectors) but otherwise you can pretty much include the entirety of a 1995 PC computer (chips and firmware, OS, etc.) inside a tiny, tiny package. The attack surface it leaves open is a different story.
- bcantrill 3y agoSo, I don't disagree in the abstract -- but in this case, the presence of these abstractions are functioning as real obstacles with respect to platform enablement, and we found it be a tremendous win to not only eliminate the abstractions entirely.[0] [0] https://www.osfc.io/2022/talks/i-have-come-to-bury-the-bios-not-to-open-it-the-need-for-holistic-systems/ https://www.osfc.io/2022/talks/i-have-come-to-bury-the-bios-...
- __d 3y agoLet's imagine a 42U rack full of Dell or HPE servers. Each one has a couple of power supplies, console ports, a video port and a GPU to drive it, USB ports for keyboard, and mouse, etc. More to the point with Oxide, there's layer upon layer of software cruft as well: various levels of proprietary firmware blobs driving the boot process, and then the management stack that underlies a commodity PC platform OS. All of that is gone, replaced with a bottom-up redesign, of both the hardware and the software, with the specific purpose of running large scale modern application loads. This goes well beyond a blade server, and is a completely different beast than a rack of DL360s.
- throwawaylinux 3y agoI know it might seem like it, but there's just not that much in it. All video, USB, "console ports" can be driven from a small chip using a couple of watts. Some do all that on their BMC SoC, even. And if your requirements include redundant power supplies, then that's what you get, and if they don't you get something without them. And the boot and runtime firmware might be complicated and have its own issues, but there are not large runtime performance overheads in that either that you can just "rewrite it all" and get a big speedup. Proof of the pudding will be in the eating I guess. Does Oxide talk about performance advantages at all, or have numbers?
- __d 3y agoA couple of watts here, and a couple of watts there, and pretty soon you're talking about real money! If it didn't matter, then there'd be no interest in blades, or OCP, and all the cloud vendors would use standard rackmount whitebox servers. At the scale of one rack, the difference is fairly small. But once you're building out cages full of racks, maybe it matters more? And it's not just power and space (although they both matter), it's the attack surface, and the ability to actually have your hardware do what you want, and the benefits of a stack that's built for purpose, not cobbled together out of off-the-shelf parts. It's not for everyone, to be sure, but I suspect that for its target market, it'll be a very successful product/family. As you say, we will see. I wish them well.
- throwawaylinux 3y ago> A couple of watts here, and a couple of watts there, and pretty soon you're talking about real money! > Not with the current state of the art. > If it didn't matter, then there'd be no interest in blades, or OCP, and all the cloud vendors would use standard rackmount whitebox servers. I'm talking about it mattering from the starting point of blades, OCP, "cloud" systems! The thread is about where the oxide niche is and what advantages it has over competition. The idea there is huge amount to be won on southbridge chipset and IO ports in large scale systems is simply not true. Quite amazing that people who don't understand this are posting in this thread as though they are experts in the matter.
- bcantrill 3y agoSo, we have found that that really IS correct, though perhaps not in the ways you might think. For example, a major problem with servers today is the accepted geometry: 1U or 2U x 19". In order to drive density, you are either doing dual socket in 2U or single socket in 1U -- but to make that geometry work you have small fans that really have to crank fans to ram air through. (And have we mentioned that the AC power supplies also need fans?). These little screamers are acoustically distasteful, but it's worse than that: they have to do much more work (and draw more power!) to move a fraction of the air because they are so small (air movement is ~cubic with respect to diameter). By changing the enclosure geometry (our sleds are 100mm high) we have much larger fans (80mm). The upshot? Our fans move SO much more air that we can run them much more slowly (we worked with our fan provider to drop the RPM at 0% PWM from 5000 RPM to 2000 RPM) -- which means our rack is not only silent by comparison, it means the power in the rack is going to compute rather than fans. So, this isn't really chasing the 0.1% -- it's chasing much bigger wins, and it's doing it by changing the power distribution (shelf of rectifiers to a 54V DC busbar vs. redundant AC power supplies in every 1U/2U), changing the cooling (the larger fans), changing the networking (we have a blindmated cabled backplane, obviating the need for operator cabling), etc. etc. These are not just multipliers on efficiency, but also on manageability: once physically installed, our rack is designed to be provisioning VMs orders of magnitude faster than the traditional manual rack/stack/cable/SW install.
- throwawaylinux 3y ago> So, we have found that that really IS correct, though perhaps not in the ways you might think. I was replying more to the idea there's all this costly legacy IO. I guess you aren't the first to try different geometries or power delivery or cabling either. There's been lots of little opencompute-type efforts and startups come and go. I'm skeptical there's a lot in it in a significant niche that does not already do these things, but you don't need to convince me. Although if you did want to you could show some comparative density numbers. What can you fit, 2048 cores, 32TB, and push 15kW through a rack?
- bcantrill 3y agoYeah: 32 AMD 7713P (64 cores/128 threads), so 2048 cores/4096 threads, 32 TB of DRAM, 1PB of NVMe, with 2x100GbE to dual, redundant 6.4Tb/s switches -- all in 15 kW. In terms of other efforts, there have certainly been some industry initiatives (OCP, Open19), but no startups that I'm aware (the smaller companies in the space have historically been integrators rather than doing their own de novo designs -- and they don't do/haven't done their own software at all); is there one in particular you're thinking of?