8 ms·
The 405b model is actually competitive against closed source frontier models. Quick comparison with GPT-4o: +----------------+-------+-------+ | M
by lelag 2y ago
The 405b model is actually competitive against closed source frontier models.
Quick comparison with GPT-4o:
+----------------+-------+-------+
| Metric | GPT-4o| Llama |
| | | 3.1 |
| | | 405B |
+----------------+-------+-------+
| MMLU | 88.7 | 88.6 |
| GPQA | 53.6 | 51.1 |
| MATH | 76.6 | 73.8 |
| HumanEval | 90.2 | 89.0 |
| MGSM | 90.5 | 91.6 |
+----------------+-------+-------+
- cchance 2y agoSuper cool, though sadly 405b will be outside most personal usage without cloud providers which sorta defeats the purpose of opensource to some extent atleast sadly, because .. nvidia's rampup of consumer VRAM is glacial
- kingsleyopara 2y agoYou might be able to get away with running a heavily quantizied 405b model using CPU inference at a blistering fast token every 5 seconds on a 7950x.
- wuschel 2y agoOK, I am curious now: What kind of hardware would I need to run such a model for a couple of users with decent performance? Where could I get a mapping of token / time vs hardware?
- angoragoats 2y agoUnsure if anyone has specific hardware benchmarks for the 405b model yet, since it's so new, but elsewhere in this thread I outlined a build that'd probably be capable of running a quantized version of Llama 3.1 405b for roughly $10k. The $10k figure is likely roughly the minimum amount of money/hardware that you'd need to run the model at acceptable speeds, as anything less requires you to compromise heavily on GPU cores (e.g. Tesla P40s also have 24GB of VRAM, for half the price or less, but are much slower than 3090s), or run on the CPU entirely, which I don't think will be viable for this model even with gobs of RAM and CPU cores, just due to its sheer size.
- bick_nyers 2y agoEnergy costs are an important factor here too. While Quadro cards are much more expensive upfront (higher $/VRAM), they are cheaper over time (lower Watts/Token). Offsetting the energy expense of a 3090/4090/5090 build via solar complicates this calculation but generally speaking can be a "reasonable" way of justifying this much hardware running in a homelab. I would be curious to see relative failure rates over time of consumer vs Quadro cards as well.
- angoragoats 2y agoAgree 100% that energy costs are important. The example system in my other post would consume somewhere around 300W at idle, 24/7, which is 219 kWh per month, and that's assuming you aren't using the machine at all. I don't have any actual figures to back this up, but my gut tells me that the fact that enterprise GPUs are an order of magnitude (at least) more expensive than, say a, 3090, means that the payback period of them has got to be pretty long. I also wonder whether setting the max power on a 3090 to a lower than default value (as I suggest in my other post) has a significant effect on the average W/token.
- bick_nyers 2y agoAgreed, but there are other costs associated with supporting 10-16x GPUs that may not necessarily happen with say 6 GPUs. Having to go from single socket (or Threadripper) to dual socket, PCIE bifurcation, PLX risers, etc. Not necessarily saying that Quadros are cheaper, just that there's more to the calculation when trying to run 405B size models at home
- angoragoats 2y agoThe system I outlined in my other post [0] has ten GPUs and does not require dual socket CPUs as far as I'm aware. It could likely scale easily to 14 GPUs as well (assuming you have sufficient power), with an x8/x8 bifurcation adapter installed in each PCIe slot. This is pushing the limits of the PCIe subsystem I'm sure, but you could also likely scale up to 28 GPUs, again assuming sufficient power, by simply bifurcating at x4/x4/x4/x4 vs x8/x8. I think it should work as-is with the components listed, but if you disagree please let me know! [0] https://news.ycombinator.com/item?id=41047689 https://news.ycombinator.com/item?id=41047689
- microtonal 2y agoYou can run the 4-bit GPTQ/AWQ quantized Llama 405B somewhat reasonably on 4x H100 or A100. You will be somewhat limited in how many tokens you can have in flight between requests and you cannot create CUDA graphs for larger batch sizes. You can run 405B well on 8x H100 and A100, either with the mixed BFloat16/FP8 checkpoint that Meta provided or GPTQ/AWQ-quantized models. Note though that the A100 does not have native support for FP8, but FP8 quantized weights can be used through the GPTQ-Marlin FP8 kernel. Here are some TGI 405B benchmarks that I did with the different quantized models: https://x.com/danieldekok/status/1815814357298577718 https://x.com/danieldekok/status/1815814357298577718 The 405B model is very useful outside direct use in inference though. E.g. for generating synthetic data for training smaller model: https://huggingface.co/blog/synthetic-data-save-costs https://huggingface.co/blog/synthetic-data-save-costs
- duchenne 2y agoMost SMBs would be able to run it. This is already a huge win for decentralized AI.
- aabhay 2y agoZoom out a bit. There’s a massive feeder ecosystem around llama. You’ll see many startups take this on and help drive down inference costs for everyone and create competitive pressure that will improve the state of the art.
- deleted 2y ago[deleted]
- gkk 2y agoIf you think of open source as a protocol through which the ecosystem of companies loosely collaborate, then it's a big deal. E.g. Groq can work on inference without a complicated negotiations with Meta. Ditto for Huggingface, and smaller startups. I agree with you on open source in the original, home tinkerer sense.
- loudmax 2y agoI agree that 405B isn't practical for home users, but I disagree that it defeats the purpose of open source. If you're building a business on inference it can be valuable to run an open model on hardware that you control, without the need to worry that OpenAI or Anthropic or whoever will make drastic changes to the model performance or pricing. Also, it allows the possibility of fine-tuning the model to your requirements. Meta believes it's in their interest to promote these businesses. I'd think of the 405B model as the equivalent to a big rig tractor trailer. It's not for home use. But also check out the benchmark improvements for the 70B and 8B models.
- monkeydust 2y agoGreat for Groq whos already hosting it but at what cost I guess.
- niutech 2y agoGroq provides a limited free tier for now: https://wow.groq.com/now-available-on-groq-the-largest-and-most-capable-openly-available-foundation-model-to-date-llama-3-1-405b/ https://wow.groq.com/now-available-on-groq-the-largest-and-m...
- stuckinhell 2y ago100% reddit is full of people trying to solder more vram
- gaogao 2y agoI've been wondering if you could just attach a chunk of vram over NVLink, since that's very roughly what FSDP is doing here anyways.
- bick_nyers 2y agoThe best NVLINK you can reasonably purchase is for the 3090, which is capped somewhere around 100 Gbit/s. This is too slow. The 3090 has about 1 TB/s memory bandwidth, and the 4090 is even faster, and the 5090 will be even faster. PCIE 5.0 x16 is 500 Gbit/s if I'm not mistaken, so using RAM is more viable an alternative in this case. Edit: 3090 has 1 TB/s, not terabits
- paxys 2y agoYou don't need a model of this scale for personal use. Llama 3.1 8B can easily run on your laptop right now. The 70B model can run on a pair of 4090s.
- api 2y agoI have the 70b model running quantized just fine on an M1 Max laptop with 64GiB unified RAM. Performance is fine and so far some Q&A tests are impressive. This is good enough for a lot of use cases... on a laptop. An expensive laptop, but hardware only gets better and cheaper over time.
- buu700 2y agoI don't have the hardware to confirm this, so I'd take it with a grain of salt, but ChatGPT tells me that a maxed out M3 MacBook Pro with 128 GB RAM should be capable of efficiently running Llama 3.1 405B, albeit with essentially no ability to multitask. (It also predicted that a MacBook Air in 2030 will be able to do the same, and that for smartphones to do the same might take around 20 years.)
- hmottestad 2y agoI’ve run the Falcon 180B on my M3 Max with 128 GB of memory. I think I ran it at 3-bit. Took a long time to load and was incredibly slow at generating text. Even if you could load the Llama 405B model it would be too slow to be of much use.
- buu700 2y agoAh, that's a shame to hear. FWIW, ChatGPT did also suggest that there was a lot of room for improvement in the MPS backend of PyTorch that would likely make it more efficient on Apple hardware in time.
- Klaus23 2y agoYou fundamentally misunderstand the bottleneck of large LLMs. It is not really possible to make gains that way. A 405B LLM has 405 billion parameters. If you run it at full "prescision", each parameter takes up 2 bytes, which means you need 810GB of memory. If it does not fit in RAM or GPU memory it will swap to disc and be unusably slow. You can run the model at reduced prescision to save memory, called quantisation, but this will degrade the quality of the response. The exact amount of degradation depends on the task, the specific model and its size. Larger models seem to suffer slightly less. 1 byte per parameter is pretty much as good as full precision. 4 bits per parameter is still good quality, 3 bits is noticeably worse and 2 bits is often bad to unusable. With 128GB of RAM, zero overhead and a 405B model, you would have to quantize to about 2.5 bits, which would noticeably degrade the response quality. There is also model pruning, which removes parameters completely, but this is much more experimental than quantisation, also degrades response quality, and I have not seen it used that widely.
- Aurornis 2y ago> sorta defeats the purpose of opensource to some extent Not in the slightest. They even have a table of cloud providers where you can host the 405B model and the associated cost to do so on their website: https://llama.meta.com/ https://llama.meta.com/ (Scroll down) "Open Source" doesn't mean "You can run this on consumer hardware". It just means that it's open source. They also released 8B and 70B models for people to use on consumer gear.
- diego_sandoval 2y agoThe fact that it takes $20k to run your own SOTA model, instead of the $2B+ that it took until yesterday, is significant.
- lostmsu 2y agoYou can probably run it on your local PC at 1 token/minute.
- bamboozled 2y agoThis nodel is not “open source”, free to use maybe.
- nomel 2y agoI really wish people would use "open weights" rather than "open source". It's precise and obvious, and leaves an accurate descriptor for actual "open source" models, where the source and methods that that generate the artifact, that is the weights, is open.
- fngjdflmdflg 2y agoAs far as I know it's not just the weights. it's everything but the dataset. So the code used to generate the weights is also open source.
- TeMPOraL 2y agoIn other words, it's everything except the one thing that actually matters.
- OrangeMusic 2y agoMaybe, but it doesn't mean it's not open source.
- TeMPOraL 2y agoThe things that don't matter are, the thing that does isn't. Together, they can hardly be called open source.
- kibibu 2y agoThe dataset is likely absolutely jam packed with copyrighted material that cannot be distributed.
- nomel 2y ago
- mi_lk 2y agoHow do you draw/generate such ascii table?
- TacticalCoder 2y agoDon't know about OP but I generate such tables using Emacs.
- deleted 2y ago[deleted]
- lelag 2y agoIn the past, I might have used a python library like asciitable to do that. This time, I just copy pasted the raw metrics I found and asked an LLM to format it as an ASCII table.
- deleted 2y ago[deleted]