5 ms·
If you didn’t have a $5,000 MacBook, could you run it on a GPU over a long period of time? Would you be able to feed it a list of prompts that maybe take all ni
by notadev 4y ago
If you didn’t have a $5,000 MacBook, could you run it on a GPU over a long period of time? Would you be able to feed it a list of prompts that maybe take all night to process and have the responses by morning?
- itake 4y agoMy m1 max was $3k and it runs the 30b model
- DennisAleynikov 4y ago30B also loads great on my M2 Pro Mac mini that's just shy of $2K This is insane
- dheera 4y agoUh, if your budget is <$5000 or <$3000 or whatever you can get a lot more GPU crunch power if you don't buy a Mac. Vanilla Ubuntu desktop with a used server-grade NVIDIA GPU with lots of GPU RAM is probably your best bang for buck for messing with large model inference right now.
- BulgarianIdiot 4y agoYou're dispensing outdated advice. Instead of saying "you'll save" and "lots of GPU RAM" start specific listing numbers and you'll soon realize your mistake. Mac has unified memory architecture that no other mainstream desktop can offer right now. This is the source of having up to 96GB VRAM on a laptop.
- dheera 4y agoOh that's cool. On a separate note I never understood why NVIDIA was always so stingy on RAM. Give us a GPU with 256GB of RAM already. It doesn't need to be that fast, it just needs to be big enough to hold a huge model. Instead they're still selling RTX4090Ti with 24GB of RAM while the free GPU I got from them at a raffle a few years ago has 32GB. WTF? I would have expected the RTX4090Ti to have 64GB minimum, being the highest-end consumer GPU.
- BulgarianIdiot 4y agoHardware lags use cases by a few years. You don't need that much RAM to game, or to mine crypto. With AI, you need all the RAM. So expect NVIDIA to offer some drastically different products in coming years. But to me, Mac's architecture is the most inevitable final solution: the GPU becomes part of the CPU, and there's no CPU RAM and GPU RAM, just RAM. Will be great for AI too.
- jandrese 4y agoMaybe we need GPUs with DIMM slots? Or maybe a second card that's all memory and connected by the bus they use for dual GPU setups? Sadly both are almost certainly too niche to exist in our current GPU marketplace.
- dheera 4y agoI feel like they would not be too niche, many AI researchers I know complain about the amount of RAM on a single GPU, especially the ones they or their labs can afford. Very often multi-GPU training is maxing out RAM on each GPU and isn't making full use of compute capacity.
- BulgarianIdiot 4y agoGPUs have too many cores, and each of them needs to read and write to RAM really fast, so any such indirection introduced through interfaces like DIMM is an obstacle to performance. I suspect the future is tightly integrated CPU+GPU+RAM, but we may STILL have DIMM slots to acts as a buffer between "fast" RAM and disk. So your swap file goes from "fast RAM" to "DIMM RAM" before it has to go to disk, basically.
- dheera 4y agoConceptually that's already the case i.e. L1 cache -> L2 cache -> RAM -> swapfile If the definition of RAM is DIMMs then it's really just a matter of merging the CPU and GPU into one unit with a massive L2 cache (or maybe we have an L3 cache) of e.g. 32GB that's able to hold a massive model without cache misses.
- hnfong 4y agoI think everyone running LLaMA on "cheap" Macs are using this: https://news.ycombinator.com/item?id=35100086 https://news.ycombinator.com/item?id=35100086 I've had success running the 7B model on a 16GB M1 Mac Mini (purchased in 2020!). I didn't bother to try but pretty sure the 13B model works as well. I suspect you could grab one of those for less than $1000 these days.