7 ms·
Big GPUs don't need big PCs
- jonahbenton 10mo agoSo glad someone did this. Have been running big gpus on egpus connected to spare laptops and thinking why not pis.
- deleted 10mo ago[deleted]
- 3eb7988a1663 10mo agoDatapoints like this really make me reconsider my daily driver. I should be running one of those $300 mini PCs at <20W. With ~flat CPU performance gains, would be fine for the next 10 years. Just remote into my beefy workstation when I actually need to do real work. Browsing the web, watching videos, even playing some games is easily within their wheelhouse.
- ekropotin 10mo agoAs experiment, I decided to try using proxmox VM with eGPU and usb bus bypassed to it, as my main PC for browsing and working on hobby projects. It’s just 1 vCPU with 4 Gb ram, and you know what? It’s more than enough for these needs. I think hardware manufactures falsely convinced us that every professional needs beefy laptop to be productive.
- reactordev 10mo agoI went with a beelink for this purpose. Works great. Keeps the desk nice and tidy while “the beasts” roar in a soundproofed closet.
- samuelknight 10mo agoSwitching from my 8-core ryzen minipc to an 8-core ryzen desktop makes my unit tests run way faster. TDP limits can tip you off to very different performance envelopes in otherwise similar spec CPUs.
- loeg 10mo agoEven if you could cool the full TDP in a micro PC, in a full size desktop you might be able to use a massive AIO radiator with fans running at very slow, very quiet speeds instead of jet turbine howl in the micro case. The quiet and ease of working in a bigger space are mostly a good tradeoff for a slightly larger form factor under a desk.
- adrian_b 10mo agoA full-size desktop computer will always be much faster for any workload that fully utilizes the CPU. However, a full-size desktop computer seldom makes sense as a personal computer, i.e. as the computer that interfaces to a human via display, keyboard and graphic pointer. For most of the activities done directly by a human, i.e. reading & editing documents, browsing Internet, watching movies and so on, a mini-PC is powerful enough. The only exception is playing games designed for big GPUs, but there are many computer users who are not gamers. In most cases the optimal setup is to use a mini-PC as your personal computer and a full-size desktop as a server on which you can launch any time-consuming tasks, e.g. compilation of big software projects, EDA/CAD simulations, testing suites etc. The desktop used as server can use Wake-on-LAN to stay powered off when not needed and wake up whenever it must run some task remotely.
- whatevaa 10mo agoNot everything supports remoting well. For example, many IDE's. Unless you run RDP, with whole graphical session on remote. Also, having to buy two computers also costs money. It makes sense to use 1 for both use cases if you have to buy the desktop anyway.
- jasonwatkinspdx 10mo agoFor just basic windows desktop stuff, a $200 NUC has been good enough for like 15 years now.
- themafia 10mo ago> I should be running one of those $300 mini PCs at <20W. Yes. They're basically laptop chips at this point. The thermals are worse but the chips are perfectly modern and can handle reasonably large workloads. I've got an 8 core Ryzen 7 with Radeon 780 Graphics and 96GB of DDR5. Outside of AAA gaming this thing is absolutely fine. The power draw is a huge win for me. It's like 6W at idle. I live remotely so grid power is somewhat unreliable and saving watts when using solar batteries extends their lifetime massively. I'm thrilled with them.
- PunchyHamster 10mo agoSlapping $300 worth of solar panels on your roof/balcony will probably get you ahead on power usage
- nottorp 10mo agoThat's why I use a M2 (not even pro) Mac Mini as a terminal and remote into other boxes when needed.
- ivanjermakov 10mo agoAnother benefit is low noise. Many consider fan noise under load to be the most important property of a workstation.
- yjftsjthsd-h 10mo agoI've been kicking this around in my head for a while. If I want to run LLMs locally, a decent GPU is really the only important thing. At that point, the question becomes, roughly, what is the cheapest computer to tack on the side of the GPU? Of course, that assumes that everything does in fact work; unlike OP I am barely in a position to understand eg. BAR problems, let alone try to fix them, so what I actually did was build a cheap-ish x86 box with a half-decent GPU and called it a day:) But it still is stuck in my brain: there must be a more efficient way to do this, especially if all you need is just enough computer to shuffle data to and from the GPU and serve that over a network connection.
- zeusk 10mo agoGet the DGX Spark computers? They’re exactly what you’re trying to build.
- Gracana 10mo agoThey’re very slow.
- geerlingguy 10mo agoThey're okay, generally, but slow for the price. You're more paying for the ConnectX-7 networking than inference performance.
- Gracana 10mo agoYeah, I wouldn’t complain if one dropped in my lap, but they’re not at the top of my list for inference hardware. Although... Is it possible to pair a fast GPU with one? Right now my inference setup for large MoE LLMs has shared experts in system memory, with KV cache and dense parts on a GPU, and a Spark would do a better job of handling the experts than my PC, if only it could talk to a fast GPU. [edit] Oof, I forgot these have only 128GB of RAM. I take it all back, I still don’t find them compelling.
- tcdent 10mo ago
- Wowfunhappy 10mo agoI really would have liked to see gaming performance, although I realize it might be difficult to find a AAA game that supports ARM. (Forcing the Pi to emulate x86 with FEX doesn't seem entirely fair.)
- 3eb7988a1663 10mo agoYou might have to thread the needle to find a game which does not bottleneck on the CPU.
- deleted 10mo ago[deleted]
- kristjansson 10mo agoReally why have the PCI/CPU artifice at all? Apple and Nvidia have the right idea: put the MPP on the same die/package as the CPU.
- bigyabai 10mo ago> put the MPP on the same die/package as the CPU. That would help in latency-constrained workloads, but I don't think it would make much of a difference for AI or most HPC applications.
- deleted 10mo ago[deleted]
- PunchyHamster 10mo agoWe need low power but high PCIE lane count CPUs for that. Just purely for shoving models from NVMe to GPU
- lostmsu 10mo agoNow compare batched training performance. Or batched inference. Of course prefill is going to be GPU bound. You only send a few thousand bytes to it, and don't really ask to return much. But after prefill is done, unless you use batched mode, you aren't really using your GPU for anything more that it's VRAM bandwidth.
- deleted 10mo ago[deleted]
- numpad0 10mo agoNot sure what was unexpected about the multi GPU part. It's very well known that most LLM frameworks including llama.cpp splits models by layers, which has sequential dependency, and so multi GPU setups are completely stalled unless there are n_gpu users/tasks running in parallel. It's also known that some GPUs are faster in "prompt processing" and some in "token generation" that combining Radeon and NVIDIA does something sometimes. Reportedly the inter-layer transfer sizes are in kilobyte ranges and PCIe x1 is plenty or something. It takes appropriate backends with "tensor parallel" mode support, which splits the neural network parallel to the direction of flow of data, which also obviously benefit substantially from good node interconnect between GPUs like PCIe x16 or NVlink/Infinity Fabric bridge cables, and/or inter-GPU DMA over PCIe(called GPU P2P or GPUdirect or some lingo like that). Absent those, I've read somewhere that people can sometimes see GPU utilization spikes walking over GPUs on nvtop-style tools. Looking for a way to break up tasks for LLMs so that there will be multiple tasks to run concurrently would be interesting, maybe like creating one "manager" and few "delegated engineers" personalities. Or simulating multiple different domains of brain such as speech center, visual cortex, language center, etc. communicating in tokens might be interesting in working around this problem.
- zozbot234 10mo ago> Looking for a way to break up tasks for LLMs so that there will be multiple tasks to run concurrently would be interesting, maybe like creating one "manager" and few "delegated engineers" personalities. This is pretty much what "agents" are for. The manager model constructs prompts and contexts that the delegated models can work on in parallel, returning results when they're done.
- nodja 10mo ago> Reportedly the inter-layer transfer sizes are in kilobyte ranges and PCIe x1 is plenty or something. Not an expert, but napkin math tells me that more often that not this will be in the order of megabytes—not kilobytes—since it scales with sequence length. Example: Qwen3 30B has a hidden state size of 5120, even if quantized to 8 bits that's 5120 bytes per token. It would pass the MB boundary with just a little over 200 tokens. Still not much of an issue when a single PCIe lane is ~2GB/s. I think device to device latency is more of an issue here, but I don't know enough to assert that with confidence.
- deleted 10mo ago[deleted]
- kgeist 10mo agoWhat about constrained decoding (with JSON schemas)? I noticed my vLLM instance is using 1 CPU 100%.
- jauntywundrkind 10mo agoPCIe 3.0 is the nice easy convenient generation where 1 lane = 1GBps. Given the overhead, thats pretty close to 10Gb ethernet speeds (low latency though). I do wonder how long the cards are going to need host systems at all. We've already seen GPUs with m.2 ssd attached! Radeon Pro SSG hails back from 2016! You still need a way to get the model on that in the first place to get work in and out, but a 1Gbe and small RISC-V chip (which Nvidia already uses formanagement cores) could suffice. Maybe even an rpi on the card. https://www.techpowerup.com/224434/amd-announces-the-radeon-pro-ssg https://www.techpowerup.com/224434/amd-announces-the-radeon-... Given the gobs of memory cards have, they probably don't even need storage; they just need big pipes. Intel had 100Gbe on their Xeon & Xeon Phi cores (10x what we saw here!) in 2016! GPUs that just plug into the switch and talk across 400Gbe or UltraEthernet or switched CXL, that run semi independently, feel so sensible, so not outlandish. https://www.servethehome.com/next-generation-interconnect-intel-omni-path-released/ https://www.servethehome.com/next-generation-interconnect-in... It's far off for now, but flash makers are also looking at radically many channel flash, which can provide absurdly high GB/s, High Bandwidth Flash. And potentially integrated some extremely parallel tensorcores on each channel. Switching from DRAM to flash for AI processing could be a colossal win for fitting large models cost effectively (& perhaps power efficiently) while still having ridiculous gobs of bandwidth. With that possible win of doing processing & filtering extremely near to the data too. https://www.tomshardware.com/tech-industry/sandisk-and-sk-hynix-join-forces-to-standardize-high-bandwidth-flash-memory-a-nand-based-alternative-to-hbm-for-ai-gpus-move-could-enable-8-16x-higher-capacity-compared-to-dram https://www.tomshardware.com/tech-industry/sandisk-and-sk-hy...
- deleted 10mo ago[deleted]
- Avlin67 10mo agotired of jeff glinglin everywhere...
- manarth 10mo agoI personally find his work and his posts interesting, and enjoy seeing them pop up on HN. If you prefer not to see his posts on the HN list pages, a practical solution is to use a browser extension (such as Stylus) to customise the HN styling to hide the posts. Here is a specific CSS style which will hide submissions from Jeff's website: tr.submission:has(td a[href="from?site=jeffgeerling.com"]), tr.submission:has(td a[href="from?site=jeffgeerling.com"]) + tr, tr.submission:has(td a[href="from?site=jeffgeerling.com"]) + tr + tr { opacity: 0.05 } In this example, I've made it almost invisible, whilst it still takes up space on the screen (to avoid confusion about the post number increasing from N to N+2). You could use { display: none } to completely hide the relevant posts. The approach can be modified to suit any origin you prefer to not come across. The limitation is that the style modification may need refactoring if HN changes the markup structure.
- mythoughtsexact 9mo agoYou're awesome, thank you. I stopped following this guy back in 2015 when he straight up forked all of my ansible roles and then published everything to Ansible Galaxy before mine were even complete, tested and ready to be published, and only for me to find that the same day they were all forked by him a new Github organization with the name of the org I had used in my roles had been registered and then squatted, it completely turned me off to his methods.
- mjh2539 10mo agoI only ever see him on HN. He's smart, kind, and talks about interesting things. Are you sure what you're feeling isn't envy?
- omneity 10mo agoI wish for a hardware + software solution to enable direct PCIe interconnect using lanes independent from the chipset/CPU. A PCIe mesh of sorts. With the right software support from say pytorch this could suddenly make old GPUs and underpowered PCs like in TFA into very attractive and competitive solutions for training and inference.
- snuxoll 10mo agoPCIe already allows DMA between peers on the bus, but, as you pointed out, the traces for the lanes have to terminate somewhere. However, it doesn't have to be the CPU (which is, of course, the PCIe root in modern systems) handling the traffic - a PCIe switch may be used to facilitate DMA between devices attached to it, if it supports routing DMA traffic directly.
- ComputerGuru 10mo agoThat’s what happened in TFA.
- omneity 10mo agoYou're right. Let me correct myself: a hobbyist-friendly hardware solution. Dolphin's PCIe switches cost more than 8 RTX 3090 on a Threadripper machine.
- ComputerGuru 10mo agoJeff forgot to mention that in his post!
- deleted 10mo ago[deleted]
- Waterluvian 10mo agoAt what point do the OEMs begin to realize they don’t have to follow the current mindset of attaching a GPU to a PC and instead sell what looks like a GPU with a PC built into it?
- nightshift1 10mo agoExactly. With the Intel-Nvidia partnership signed this September, I expect to see some high-performance single-board computers being released very soon. I don't think the atx form-factor will survive another 30 years.
- bostik 10mo agoOne should also remember that NVidia does have organisational experience on designing and building CPUs[0]. They were a pretty big deal back in ~2010, and I have to admit I didn't know that Tegra was powering Nintendo Switch. 0: https://en.wikipedia.org/wiki/Tegra https://en.wikipedia.org/wiki/Tegra
- goku12 10mo agoI had a Xolo Tegra Note 7 tablet (marketed in the US as EVGA Tegra Note 7) in around 2013. I preordered it as far as I remember. It had a Tegra 4 SoC with quad core Cortex A15 CPU and a 72 core GeForce GPU. Nvidia used to claim that it is the fastest SoC for mobile devices at the time. To this day, it's the best mobile/Android device I ever owned. I don't know if it was the fastest, but it certainly was the best performing one I ever had. UI interactions were smooth, apps were fast on it, screen was bright, touch was perfect and still had long enough battery backup. The device felt very thin and light, but sturdy at the same time. It had a pleasant matte finish and a magnetic cover that lasted as long as the device did. It spolied the feel of later tablets for me. It had only 1 GB RAM. We have much more powerful SoCs today. But nothing ever felt that smooth (iPhone is not considered). I don't know why it was so. Perhaps Android was light enough for it back then. Or it may have had a very good selection and integration of subcomponents. I was very disappointed when Nvidia discontinued the Tegra SoC family and tablets.
- deleted 10mo ago
- pjmlp 10mo agoOf course, just go to any computer store where most gamer setups on affordable bugets go with the combo "beefy GPU + an i5", instead of using an i7 or i9 Intel CPUs.
- moebrowne 10mo agoI'd be interested to see if workloads like Folding@home could be efficiently run this way. I don't think they need a lot of bandwidth.
- haritha-j 10mo agoI currently have a £500 laptop hooked up to an egpu box with a £700 gpu. It's not a bad setup.
- deleted 10mo ago[deleted]
- nailherwithrust 10mo ago[flagged]
- yoan9224 10mo agoThe most interesting takeaway for me is that PCIe bandwidth really doesn't bottleneck LLM inference for single-user workloads. You're essentially just shuttling the model weights once, then the GPU churns through tokens using its own VRAM. This is huge for home lab setups. You can run a Pi 5 with a high-end GPU via external enclosure and get 90% of the performance of a full workstation at a fraction of the power draw and cost. The multi-GPU results make sense too - without tensor parallelism, you're just pipeline parallelism across layers, which is inherently sequential. The GPUs are literally sitting idle waiting for the previous layer's output. Exo and similar frameworks are trying to solve this but it's still early days. For anyone considering this: watch out for ResizeBAR requirements. Some older boards won't work at all without it.