18 ms·
Llama 3.1
- TechDebtDevin 2y agoNice, someone donate me a few 4090s :(
- foxhop 2y agoYour going to need a lot more than a few, 800G VRAM needed
- TechDebtDevin 2y agoOof.
- glitchc 2y agoChrist!!
- lolinder 2y agoQuantized to 4 bits you'll only need ~200GB! 5 4090s should cover it.
- angoragoats 2y agoYou'll probably need 9 or more. 4090s have 24GB each.
- woodson 2y agoI wonder if AutoAWQ works out of the box, given no architectural changes (?). That would be most straightforward together with vLLM for serving.
- pat2man 2y agoTwo 128gb Mac studios networked via thunderbolt 4?
- Teknomancer 2y agoThis is actually a promising endeavor. Id love to see someone try that.
- angoragoats 2y agoThere's already at least one project that attempts this: https://github.com/exo-explore/exo https://github.com/exo-explore/exo
- downvotetruth 2y agoIf an implementation had NVidia's Heterogeneous Memory Management implemented, then 192 GB RAM DDR5 + GPU VRAM would seem to be close.
- AaronFriel 2y agoIf previous quantization results hold up, fp8 will have nearly identical performance while using 405GiB for weights, but the KV cache size will still be significant. Too bad, too, I don't think my PC will fit 20 4090s (480GiB).
- knicholes 2y agoI've got a motherboard that will support 8!
- Zambyte 2y ago40,320 4090s?? What witchcraft is this?! :D
- sebastiennight 2y agoAll the more impressive when you realize that Groq's infrastructure (based on LPUs) was built using only 6!
- beeboobaa3 2y agohow is this even useful? no one can run it.
- jermaustin1 2y agoYou don't use the 405B parameter model at home. I have a lot of luck with 8B and 13B models on a single 3090. You can quantize them down (is that the term) which lowers precision and memory use, but still very usable... most of the time. If you are running a commercial service that uses AI, you buy a few dozen A100s, spend a half million, and you are good for a while. If you are running a commercial inferencing service, you spend tens of millions or get a cloud sponsor.
- beeboobaa3 2y agoI can't expect all my users to have 3090s and if we're talking about spending millions there are better things to invest in than a stack of GPUs that will be obsolete in a year or three.
- jermaustin1 2y agoNo, but if you are thinking about edge compute for LLMs, you quantize. Models are getting more efficient, and there are plenty of SLMs and smaller LLMs (like phi-2 or phi-3) that are plenty capable even on a tiny arm device like the current range of RPi "clones". I have done experiments with 7B Llama3 Q8 models on a M3 MBP. They run faster than I can read, and only occasionally fall off the rails. 3B Phi-3 mini is almost instantaneous in simple responses on my MBP. When I want longer context windows, I use a hosted service somewhere else, but if I only need 8000 tokens (99% of the time that is MORE than I need), any of my computers from the last 3 years are working just fine for it.
- loudmax 2y agoIf you want to run the 405B model without spending thousands of dollars on dedicated hardware, you rent compute from a datacenter. Meta lists AWS, Google and Microsoft among others as cloud partners. But also check out the 8B and 70B Llama-3.1 models which show improved benchmarks over the Llama-3 models released in April.
- whalesalad 2y agofollow the trail of tears to my credit card
- lawlessone 2y agomaybe someone will figure out some ways to prune/ quantize it a huge amount ;-; edit: If the AI bubble pops we will be swimming in GPUs... but no new models.
- Sakthimm 2y agoThis is absurd. We have crossed the point of no return, llms will forever be in our lives in one form or another, just like internet, especially with the release of these open model weights. There is no bubble, only way forward is better, efficient llms, everywhere.
- tymscar 2y agoYou seem to not understand what a bubble popping is. Yes we have the internet around, that doesn’t mean the dot com bubble didn’t pop…
- yard2010 2y agoThis bubble collapsing along with most blockchains going all in with proof of stake rather than proof of work is myself and every other gamer wet dream.
- daft_pink 2y agoWhat kind of machine do I need to run 405B local?
- 93po 2y agoaccording to another comment, ~10x 4090 video cards.
- Teknomancer 2y agoThat was the punchline of a joke.
- 93po 2y agolol thanks, i know nothing about the hardware side of things for this stuff
- daft_pink 2y agothanks. hoping the Nvidia 50 series offers some more VRAM.
- monkmartinez 2y agoYou can't. Sorry. Unless... You have a couple hundred $k sitting around collecting dust... then all you need is a DGX or HGX level of vRAM, the power to run it, the power to keep it cool, and place for it to sit.
- causal 2y agoYou could run a 4bit quant for about $10k I'm guessing. 10x3090s would do.
- angoragoats 2y agoYou can build a machine that will run the 405b model for much, much less, if you're willing to accept the following caveats: * You'll be running a Q5(ish) quantized model, not the full model * You're OK with buying used hardware * You have two separate 120v circuits available to plug it into (I assume you're in the US), or alternatively a single 240v dryer/oven/RV-style plug. The build would look something like (approximate secondary market prices in parentheses): * Asrock ROMED8-2T motherboard ($700) * A used Epyc Rome CPU ($300-$1000 depending on how many cores you want) * 256GB of DDR4, 8x 32GB modules ($550) * nvme boot drive ($100) * Ten RTX 3090 cards ($700 each, $7000 total) * Two 1500 watt power supplies. One will power the mobo and four GPUs, and the other will power the remaining six GPUs ($500 total) * An open frame case, the kind made for crypto miners ($100?) * PCIe splitters, cables, screws, fans, other misc parts ($500) Total is about $10k, give or take. You'll be limiting the GPUs (using `nvidia-smi` or similar) to run at 200-225W each, which drastically reduces their top-end power draw for a minimal drop in performance. Plug each power supply into a different AC circuit, or use a dual 120V adapter with a 240V outlet to effectively accomplish the same thing. When actively running inference you'll likely be pulling ~2500-2800W from the wall, but at idle, the whole system should use about a tenth of that. It will heat up the room it's in, especially if you use it frequently, but since it's in an open frame case there are lots of options for cooling. I realize that this setup is still out of the reach of the "average Joe" but for a dedicated (high-end) hobbyist or someone who wants to build a business, this is a surprisingly reasonable cost. Edit: the other cool thing is that if you use fast DDR4 and populate all 8 RAM slots as I recommend above, the memory bandwidth of this system is competitive with that of Apple silicon -- 204.8GB/sec, with DDR4-3200. Combined with a 32+ core Epyc, you could experiment with running many models completely on the CPU, though Lllama 405b will probably still be excruciatingly slow.
- casper14 2y agoDamn 405b params
- ThrowawayTestr 2y agoAre there any other models with free unlimited use like chatgpt?
- hubraumhugo 2y agoI wrote about this when llama-3 came out, and this launch confirms it: Meta's goal from the start was to target OpenAI and the other proprietary model players with a "scorched earth" approach by releasing powerful open models to disrupt the competitive landscape. Meta can likely outspend any other AI lab on compute and talent: - OpenAI makes an estimated revenue of $2B and is likely unprofitable. Meta generated a revenue of $134B and profits of $39B in 2023. - Meta's compute resources likely outrank OpenAI by now. - Open source likely attracts better talent and researchers. - One possible outcome could be the acquisition of OpenAI by Microsoft to catch up with Meta. The big winners of this: devs and AI product startups
- foolswisdom 2y agoIt could definitely be seen as part of that strategy, but do you mind elaborating why you think "this launch confirms it"?
- adam_arthur 2y agoIt's pretty clear the base model is a race to the bottom on pricing. There is no defensible moat unless a player truly develops some secret sauce on training. As of now seems that the most meaningful techniques are already widely known and understood. The money will be made on compute and on applications of the base model (that are sufficiently novel/differentiated). Investors will lose big on OpenAI and competitors (outside of greater fool approach)
- lolinder 2y ago> There is no defensible moat unless a player truly develops some secret sauce on training. This is why Altman has gone all out pushing for regulation and playing up safety concerns while simultaneously pushing out the people in his company that actually deeply worry about safety. Altman doesn't care about safety, he just wants governments to build him a moat that doesn't naturally exist.
- changoplatanero 2y ago> Open source likely attracts better talent and researchers I work at OpenAI and used to work at meta. Almost every person from meta that I know has asked me for a referral to OpenAI. I don’t know anyone who left OpenAI to go to meta.
- primaprashant 2y agoThe resources for link to model card[1], research paper, and Prompt Guard Tutorial[2] on the page doesn't exist yet [1]: https://github.com/meta-llama/llama-models/blob/main/models/llama3_1/MODEL_CARD.md https://github.com/meta-llama/llama-models/blob/main/models/... [2]: https://github.com/meta-llama/llama-recipes/blob/main/recipes/responsible_ai/prompt_guard/Prompt%20Guard%20Tutorial.ipynb https://github.com/meta-llama/llama-recipes/blob/main/recipe...
- meetpateltech 2y agoOpen Source AI Is the Path Forward - Mark Zuckerberg https://about.fb.com/news/2024/07/open-source-ai-is-the-path-forward/ https://about.fb.com/news/2024/07/open-source-ai-is-the-path...
- gkfasdfasdf 2y agoMeta the new "Open" AI?
- para_parolu 2y agoUntil they make model much better than competitors to actually start capitalizing on it
- ninjin 2y agoSo are they actually making the models open now or are they staying the course with "kind of open" as they have done for LLaMA 1, 2, and 3 [1]? [1]: https://opensource.org/blog/metas-llama-2-license-is-not-open-source https://opensource.org/blog/metas-llama-2-license-is-not-ope... As I have stated time and again, it is perfectly fine for them to slap on whatever license they see fit as it is their work. But it would be nice if they used appropriate terms so as not to disrupt the discourse further than they have already done. I have written several walls of text why I as a researcher find Facebook's behaviour problematic so I will fall back on an old link [2] this time rather than writing it all over again. [2]: https://news.ycombinator.com/item?id=38427832 https://news.ycombinator.com/item?id=38427832
- Zambyte 2y ago> it is perfectly fine for them to slap on whatever license they see fit as it is their work. Is it? Has there been a ruling on the enforceability of the license they attach to their models yet? Just because you say what you release can only be used for certain things doesn't actually mean what you say means anything.
- moffkalast 2y ago> specifically, it puts restrictions on commercial use for some users (paragraph 2) and also restricts the use of the model and software for certain purposes (the Acceptable Use Policy) It's "a Google and Apple can't use this model in production" clause that frankly we can all be relatively okay with.
- albert_e 2y agothis "Model Card" github link on [https://llama.meta.com/docs/overview/ https://llama.meta.com/docs/overview/] seems broken? https://github.com/meta-llama/llama-models/blob/main/models/llama3_1/MODEL_CARD.md https://github.com/meta-llama/llama-models/blob/main/models/...
- Vagantem 2y agoAs someone who just started generating AI landing pages for Dropory, this is music to my ears
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- denz88 2y agoI'm glad to see the nice incremental gains on the benchmarks for the 8B and 70B models as well.
- loudmax 2y agoSome of those benchmarks show quite significant gains. Going from Llama-3 to Llama-3.1, MMLU scores for 8B are up from 65.3 to 73.0, and 70B are up from 80.9 to 86.0. These scores should always be taken with a grain of salt, but this is encouraging. 405B is hopelessly out of reach for running in a homelab without spending thousands of dollars. For most people wanting to try out the 405B model, the best option is to rent compute from a datacenter. Looking forward to seeing what it can accomplish.
- sroussey 2y agoHow much can you quantize that down to run on a Mac Studio with 192GB? Is it possible? Feels like it would have to be 2bit…
- Davidzheng 2y agoLess than 2bit i think. There's this IQ2 quant that fits
- deleted 2y ago[deleted]
- Atreiden 2y agoIs there a way to run this in AWS? Seems like the biggest GPU node they have is the p5.48xlarge @ 640GB (8xH100s). Routing between multiple nodes would be too slow unless there's an InfiniBand fabric you can leverage. Interested to know if anyone else is exploring this.
- Tiberium 2y agoAWS has a separate service for running LLMs called Amazon Bedrock, it shouldn't take long for them to add 3.1 since they have 3 and 2 already.
- tpm 2y agofp8 quantization should work if that's acceptable?
- woodson 2y agoYou can run multi-node with tensor parallel plus pipeline parallel inference, e.g. with vLLM (https://docs.vllm.ai/en/latest/serving/distributed_serving.html https://docs.vllm.ai/en/latest/serving/distributed_serving.h...).
- TheAceOfHearts 2y agoDoes anyone know why they haven't released any 30B-ish param models? I was expecting that to happen with this release and have been disappointed once more. They also skipped doing a 30B-ish param model for llama2 despite claiming to have trained one.
- nickpsecurity 2y agoMaybe they think more people will just use quantized versions of 70B.
- prvc 2y agoWhy should they?
- TheAceOfHearts 2y agoUnless I'm misremembering, they announced it at one point. It's just giving people more options.
- prvc 2y agoThat was for version 2, not 3 or 3.1, if I recall correctly.
- michaelt 2y agoI suspect 30B models are in a weird spot, too big for widespread home use, too small for cutting edge performance. For home users 7B models (which can fit on an 8GB GPU) and 13B models (which can fit on a 16GB GPU) are in far more demand. If you're a researcher, you want a 70B model to get the best performance, and so your benchmarks are comparable to everyone else.
- drdaeman 2y agoI thought home use is whatever fits in 24GB (a single 3090 GPU, which is pretty affordable), not 8 or 16. 30B models fit.
- jcmp 2y ago"Meta AI isn't available yet in your country" Hi from europe :/
- sva_ 2y agoYou can load the page using a VPN and then turn off the VPN and the page will still work.
- sunaookami 2y agoYou can't sign in though, that worked before. Seems like they also check from which country your Facbook/Instagram account is. You can't create images without an account sadly.
- lawlessone 2y agoSomeone will torrent it soon enough i'm sure.
- WinstonSmith84 2y agoI changed my Facebook country (to Canada), using also a VPN to Canada, but that didn't help. That used to work before somehow.
- monkmartinez 2y agoWhy are (some) Europeans surprised when they are not included in tech product débuts? My lay understanding could best be described as; EU law is incredibly business unfriendly and takes a heroic effort in time and money to implement the myriad of requirements therein. Am I wrong?
- w4 2y ago> Why are (some) Europeans surprised when they are not included in tech product débuts? We had a brief, abnormal, and special moment in time after the crypto wars ended in the mid-2000s where software products were truly global, and the internet was more or less unregulated and completely open (at least in most of the world). Sadly it seems that this era has come to a close, and people have not yet updated their understanding of the world to account for that fact. People are also not great at thinking through the second order effects of the policies they advocate for (e.g. the GDPR), and are often surprised by the results.
- yinser 2y agoThe race to the bottom for pricing continues.
- netsec_burn 2y agoToday appears to be the day you can run an LLM that is competitive with GPT-4o at home with the right hardware. Incredible for progress and advancement of the technology. Statement from Mark: https://about.fb.com/news/2024/07/open-source-ai-is-the-path-forward/ https://about.fb.com/news/2024/07/open-source-ai-is-the-path...
- lolinder 2y ago> at home with the right hardware Where the right hardware is 10x4090s even at 4 bits quantization. I'm hoping we'll see these models get smaller, but the GPT-4-competitive one isn't really accessible for home use yet. Still amazing that it's available at all, of course!
- petercooper 2y agoIt's hardly cheap starting at about $10k of hardware, but another potential option appears to be using Exo to spread the model across a few MBPs or Mac Studios: https://x.com/exolabs_/status/1814913116704288870 https://x.com/exolabs_/status/1814913116704288870
- niutech 2y agoOr maybe using Distributed Llama? https://github.com/b4rtaz/distributed-llama https://github.com/b4rtaz/distributed-llama
- dunefox 2y agoIt's not really competitive though, is it? I tested it and 4o is just better.
- dunefox 2y agoDisclaimer: I tested llama3-8B, 3.1 might even as a small model be better, but I so far I have not seen a single small model approach 4o, ime.
- deleted 2y ago
- AaronFriel 2y agoIs there pricing available on any of these vendors? Open source models are very exciting for self hosting, but the per-token hosted inference pricing hasn't been competitive with OpenAI and Anthropic, at least for a given tier of quality. (E.g.: Llama 3 70B costing between $1 and $10 per million tokens on various platforms, but Claude Sonnet 3.5 is $3 per million.)
- handzhiev 2y agoLlama 3 is 0.59/0.79 on Groq. Still no price for 3.1
- primaprashant 2y agoI have found Claude 3.5 Sonnet really good for coding tasks along with the artifacts feature and seems like it's still the king on the coding benchmarks
- cubefox 2y agoI have found it to be better than GPT-4o at math too, despite the latter being better at several math benchmarks.
- wfme 2y agoMy experience reflects this too. My hunch is that GPT-4o was trained to game the benchmarks rather than output higher quality content. In theory the benchmarks should be a pretty close proxy for quality, but that doesn't match my experience at all.
- margorczynski 2y agoA problem with a lot of benchmarks is that they are out in the open so the model basically trains to game them instead of actually acquiring knowledge that would let it solve it. Probably private benchmarks that are not in the training set of these models should give better estimates about their general performance.
- Davidzheng 2y agoI personally disagree. But i haven't used sonnet that much
- cubefox 2y agoI asked both whether the product of two odds (odds=(probability/(1-probability)) can itself be interpreted as an odds, and if so, which. Neither could solve the problem completely, but Claude 3.5 Sonnet at least helped me to find the answer after a while. I assume the questions in math benchmarks are different.
- 2y ago
- diimdeep 2y agoThis 405B seriously need quantization solution like 1.625 bpw ternary packing for BitNet b1.58 https://github.com/ggerganov/llama.cpp/pull/8151 https://github.com/ggerganov/llama.cpp/pull/8151
- kromem 2y agoIn general this needs to be done across the board. The perplexity per parameter is higher and the delta grows as it scales. Not per bit, but per parameter. Why this is happening really needs more attention and more consideration for pretrained model development right now. A sleeping giant of a difference in a space where even marginal gains make headlines.
- lelag 2y agoThe 405b model is actually competitive against closed source frontier models. Quick comparison with GPT-4o: +----------------+-------+-------+ | Metric | GPT-4o| Llama | | | | 3.1 | | | | 405B | +----------------+-------+-------+ | MMLU | 88.7 | 88.6 | | GPQA | 53.6 | 51.1 | | MATH | 76.6 | 73.8 | | HumanEval | 90.2 | 89.0 | | MGSM | 90.5 | 91.6 | +----------------+-------+-------+
- cchance 2y agoSuper cool, though sadly 405b will be outside most personal usage without cloud providers which sorta defeats the purpose of opensource to some extent atleast sadly, because .. nvidia's rampup of consumer VRAM is glacial
- kingsleyopara 2y agoYou might be able to get away with running a heavily quantizied 405b model using CPU inference at a blistering fast token every 5 seconds on a 7950x.
- wuschel 2y agoOK, I am curious now: What kind of hardware would I need to run such a model for a couple of users with decent performance? Where could I get a mapping of token / time vs hardware?
- angoragoats 2y agoUnsure if anyone has specific hardware benchmarks for the 405b model yet, since it's so new, but elsewhere in this thread I outlined a build that'd probably be capable of running a quantized version of Llama 3.1 405b for roughly $10k. The $10k figure is likely roughly the minimum amount of money/hardware that you'd need to run the model at acceptable speeds, as anything less requires you to compromise heavily on GPU cores (e.g. Tesla P40s also have 24GB of VRAM, for half the price or less, but are much slower than 3090s), or run on the CPU entirely, which I don't think will be viable for this model even with gobs of RAM and CPU cores, just due to its sheer size.
- sfblah 2y agoIs there an actual open-source community around this in the spirit of other ones where people outside meta can somehow "contribute" to it? If I wanted to "work on" this somehow, what would I do?
- sangnoir 2y agoThere are a bunch of downstream fine-tuned and/or quantized models where people collaborate and share their recipes. In terms of contributing to Llama itself - I suspect Meta wants (or needs) code contributions at this time.
- sfblah 2y agoCan you give me a tip of where to look? I'm interested in participating.
- sangnoir 2y agoYou'll probably find interesting threads and links at https://old.reddit.com/r/LocalLLaMa https://old.reddit.com/r/LocalLLaMa
- sebastiennight 2y agoDid you mean, Meta does not want or need code contributions? It would seem to make more sense.
- sangnoir 2y agoYes - that's ehat I meant, but mangled it it while editing.
- foundval 2y agoYou can chat with these new models at ultra-low latency at groq.com. 8B and 70B API access is available at console.groq.com. 405B API access for select customers only – GA and 3rd party speed benchmarks soon. If you want to learn more, there is a writeup at https://wow.groq.com/now-available-on-groq-the-largest-and-most-capable-openly-available-foundation-model-to-date-llama-3-1-405b/ https://wow.groq.com/now-available-on-groq-the-largest-and-m.... (disclaimer, I am a Groq employee)
- sagz 2y ago405B is already being served on WhatsApp! https://ibb.co/kQ2tKX5 https://ibb.co/kQ2tKX5
- Workaccount2 2y agoHow do you get that option?
- e12e 2y agoAnd available via poe: https://poe.com/s/LCAyUbAgUx8UcVMhM3Re https://poe.com/s/LCAyUbAgUx8UcVMhM3Re
- geepytee 2y agoWe also added Llama 3.1 405B to our VSCode copilot extension for anyone to try coding with it. Free trial gets you 50 messages, no credit card required - https://double.bot https://double.bot (disclaimer, I am the co-founder)
- noble-lombax 2y agowould be great if there was a page showing benchmarks compared to other auto completion tools
- quotemstr 2y agoGroq's TSP architecture is one of the weirder and more wonderful ISAs I've seen lately. The choice of SRAM in fascinating. Are you guys planning on publishing anything about how you bridged the gap between your order-hundreds-megabytes SRAM TSP main memory and multi-TB model sizes?
- kingsleyopara 2y agoThe biggest win here has to be the context length increase to 128k from 8k tokens. Till now my understanding is there hasn't been any open models anywhere close to that.
- HanClinto 2y agoIt is notable, but it's not alone. Mistral NeMo just released last week with a 128k context window: https://news.ycombinator.com/item?id=40996058 https://news.ycombinator.com/item?id=40996058
- kingsleyopara 2y agoThanks! Not sure how I missed that :)
- HanClinto 2y agoIt's easy to miss things. Trying to keep up with the latest in AI news is like drinking from the firehose -- it's never-ending.
- cpursley 2y agoPhi 3
- sagz 2y agoThe 405B model is already being served on WhatsApp: https://ibb.co/kQ2tKX5 https://ibb.co/kQ2tKX5
- tarasglek 2y agoIs this official? How does one use this. I'm a very newbie whatsup so sorry for dumb q
- chown 2y agoWow! The benchmarks are truly impressive, showing significant improvements across almost all categories. It's fascinating to see how rapidly this field is evolving. If someone had told me last year that Meta would be leading the charge in open-source models, I probably wouldn't have believed them. Yet here we are, witnessing Meta's substantial contributions to AI research and democratization. On a related note, for those interested in experimenting with large language models locally, I've been working on an app called Msty [1]. It allows you to run models like this with just one click and features a clean, functional interface. Just added support for both 8B and 70B. Still in development, but I'd appreciate any feedback. [1]: https://msty.app https://msty.app
- downvotetruth 2y agoTried using msty today and it refused to open and demanded an upgrade from 0.9 - remotely breaking a local app that had been working is unacceptable. Good luck retaining users.
- sagz 2y agoHi! Love Msty Can you add GCP Vertex AI API support? Then one key would enable Claude, Llama herd, Gemini, Gemma etc
- d13 2y agoI love Msty too. Could you please add a feature to allow adding any arbitrary inference endpoint?
- ChrisArchitect 2y agoRelated: Open Source AI Is the Path Forward https://about.fb.com/news/2024/07/open-source-ai-is-the-path-forward/ https://about.fb.com/news/2024/07/open-source-ai-is-the-path... (https://news.ycombinator.com/item?id=41046773 https://news.ycombinator.com/item?id=41046773)
- unraveller 2y agoWhat are the substantial changes from 3.0 to 3.1 (70B) in terms of training approach? They don't seem to say how the training data differed just that both were 15T. I gather 3.0 was just a preview run and 3.1 was distilled down from the 405B somehow.
- thntk 2y agoCorrect me if I'm wrong, my impression is that 3.1 is a better fine-tuned variant of base 3.0 with extensive use of synthetic data.
- Workaccount2 2y ago@dang why was this removed/filtered from the front page?
- nomel 2y agoI see a few cloud hosting providers for it on the front page. I wonder if it's being gamed.
- dado3212 2y ago> We use synthetic data generation to produce the vast majority of our SFT examples, iterating multiple times to produce higher and higher quality synthetic data across all capabilities. Additionally, we invest in multiple data processing techniques to filter this synthetic data to the highest quality. This enables us to scale the amount of fine-tuning data across capabilities. [0] Have other major models explicitly communicated that they're trained on synthetic data? [0]. https://ai.meta.com/blog/meta-llama-3-1/ https://ai.meta.com/blog/meta-llama-3-1/
- tommy_axle 2y agoIt's in the <7B club, but Phi has always had a good dose of synthetic data https://huggingface.co/microsoft/Phi-3-mini-4k-instruct https://huggingface.co/microsoft/Phi-3-mini-4k-instruct
- usaar333 2y agoTechnically this is post training. This has been standard for a long time now - I think InstructGPT (gpt 3.5 base) was the last that used only human feedback (RLHF)
- dang 2y agoRelated ongoing thread: Open source AI is the path forward - https://news.ycombinator.com/item?id=41046773 https://news.ycombinator.com/item?id=41046773 - July 2024 (278 comments)
- ajhai 2y agoYou can already run these models locally with Ollama (ollama run llama3.1:latest) along with at places like huggingface, groq etc. If you want a playground to test this model locally or want to quickly build some applications with it, you can try LLMStack (https://github.com/trypromptly/LLMStack https://github.com/trypromptly/LLMStack). I wrote last week about how to configure and use Ollama with LLMStack at https://docs.trypromptly.com/guides/using-llama3-with-ollama https://docs.trypromptly.com/guides/using-llama3-with-ollama. Disclaimer: I'm the maintainer of LLMStack
- jxy 2y agoYou are a maintainer of a software that depends on ollama, so you should know that ollama depends on llama.cpp. And as of now, llama.cpp doesn't support the new ROPE: https://github.com/ggerganov/llama.cpp/issues/8650 https://github.com/ggerganov/llama.cpp/issues/8650, and all ollama can do is wait for llama.cpp: https://github.com/ollama/ollama/issues/5881 https://github.com/ollama/ollama/issues/5881
- ajhai 2y agoI've tested Q4 on M1 and it works though the quality may not likely be the same as you'd expect as others have pointed out on the issue.
- zhanghsfz 2y agoWe supported Llama 3.1 405B model on our distributed GPU network at Hyperbolic Labs! Come and use the API for FREE at https://app.hyperbolic.xyz/models https://app.hyperbolic.xyz/models Let us know if you have other needs!
- zhanghsfz 2y agoWe supported Llama 3.1 405B model on our distributed GPU network at Hyperbolic Labs! Come and use the API for FREE at https://app.hyperbolic.xyz/models https://app.hyperbolic.xyz/models Would love to hear your feedback!
- breadsniffer 2y agoI tried it, and it's good but I feel like the synthetic data used for training 3.1 does not hold up to gpt4o prob using human-curated data.
- IceHegel 2y agoWill 405b run on 8x H100s? Will it need to be quantized?
- bddppq 2y agoyep with <= 8bit (int8/fp8) quantization
- bick_nyers 2y agoI'm curious what techniques they used to distill the 405B model down to 70B and 8B. I gave the paper they released a quick skim but couldn't find any details.
- rcarmo 2y agoWorking great in ollama: https://mastodon.social/@rcarmo/112837520236956526 https://mastodon.social/@rcarmo/112837520236956526
- jxy 2y agohttps://github.com/ollama/ollama/issues/5881 https://github.com/ollama/ollama/issues/5881 https://github.com/ggerganov/llama.cpp/issues/8650 https://github.com/ggerganov/llama.cpp/issues/8650
- rcarmo 2y agoStill works fine for me. Latest ollama, running on NVIDIA.
- raminf 2y agoFWIW, 405B not working with Ollama on a Mac M3-pro Max with 128GB RAM. Times out.
- pbmonster 2y agoDid you get a 2 bit quant? You need to chain several Mac Studios via Exo to get enough memory for a useful quant to work.
- stiltzkin 2y agoWhatsApp now uses 70B too if you want to test it.
- zone411 2y agoI've just finished running my NYT Connections benchmark on all three Llama 3.1 models. The 8B and 70B models improve on Llama 3 (12.3 -> 14.0, 24.0 -> 26.4), and the 405B model is near GPT-4o, GPT-4 turbo, Claude 3.5 Sonnet, and Claude 3 Opus at the top of the leaderboard. GPT-4o 30.7 GPT-4 turbo (2024-04-09) 29.7 Llama 3.1 405B Instruct 29.5 Claude 3.5 Sonnet 27.9 Claude 3 Opus 27.3 Llama 3.1 70B Instruct 26.4 Gemini Pro 1.5 0514 22.3 Gemma 2 27B Instruct 21.2 Mistral Large 17.7 Gemma 2 9B Instruct 16.3 Qwen 2 Instruct 72B 15.6 Gemini 1.5 Flash 15.3 GPT-4o mini 14.3 Llama 3.1 8B Instruct 14.0 DeepSeek-V2 Chat 236B (0628) 13.4 Nemotron-4 340B 12.7 Mixtral-8x22B Instruct 12.2 Yi Large 12.1 Command R Plus 11.1 Mistral Small 9.3 Reka Core-20240501 9.1 GLM-4 9.0 Qwen 1.5 Chat 32B 8.7 Phi-3 Small 8k 8.4 DBRX 8.0
- henryaj 2y agoI love Connections! Can you tell us more about your benchmark?
- kristianp 2y agoHas anyone got a comparison of the performance of Llama 3.1 8B and the recent GPT-4o-mini?
- ofermend 2y agoI'm excited to try it with RAG and see how it performs (the 405B model)
- cpursley 2y agoWhat's your RAG approach? Dump everything into the model, chunk text and retrieve via vector store or something else?
- htk 2y agoVery insteresting! Running the 70B version on ollama on a mac and it's great. I asked to "turn off the guidelines" and it did, then I asked to turn off the disclaimers, after that I asked for a list of possible "commands to reduce potencial biases from the engineers" and it complied giving me an interesting list.
- CGamesPlay 2y agoThe LMSys Overall leaderboard <https://chat.lmsys.org/?leaderboard https://chat.lmsys.org/?leaderboard> can tell us a bit more about how these models will perform in real life, rather than in a benchmark context. By comparing the ELO score against the MMLU benchmark scores, we can see models which outperform / underperform based on their benchmark scores relative to other models. A low score here indicates that the model is more optimized for the benchmark, while a higher score indicates it's more optimized for real-world examples. Using that, we can make some inferences about the training data used, and then extrapolate how future models might perform. Here's a chart: <https://docs.getgrist.com/gV2DtvizWtG7/LLMs/p/5?embed=true https://docs.getgrist.com/gV2DtvizWtG7/LLMs/p/5?embed=true> Examples: OpenAI's GPT 4o-mini is second only to 4o on LMSys Overall, but is 6.7 points behind 4o on MMLU. It's "punching above its weight" in real-world contexts. The Gemma series (9B and 27B) are similar, both beating the mean in terms of ELO per MMLU point. Microsoft's Phi series are all below the mean, meaning they have strong MMLU scores but aren't preferred in real-world contexts. Llama 3 8B previously did substantially better than the mean on LMSys Overall, so hopefully Llama 3.1 8B will be even better! The 70B variant was interestingly right on the mean. Hopefully the 430B variant won't fall below!
- sujay1844 2y agoThese days, lmsys elo is the only thing I trust. The other benchmark scores mean nothing to me at this point
- __jl__ 2y agoI disagree. Not saying the other benchmarks are better. It just depends on your use case and application. For my use of the chat interface, I don't think lmsys is very useful. lmsys mainly evaluates relatively simple, low token count questions. Most (if not all) are single prompts, not conversations. The small models do well in this context. If that is what you are looking for, great. However, it does not test longer conversations with high token counts. Just saying that all benchmarks, including lmsys, have issues and are focused on specific use cases.
- Lockal 2y ago
- ofou 2y agoLlama 3 Training System 19.2 exaFLOPS _____ / \ Cluster 1 Cluster 2 / \ 9.6 exaFLOPS 9.6 exaFLOPS / \ _______ _______ / ___ \ / \ / \ ,----' / \`. `-' 24000 `--' 24000 `----. ( _/ __) GPUs GPUs ) `---'( / ) 400+ TFLOPS 400+ TFLOPS ,' \ ( / per GPU per GPU ,' \ \/ ,' \ \ TOTAL SYSTEM ,' \ \ 19,200,000 TFLOPS ,' \ \ 19.2 exaFLOPS ,' \___\ ,' `----------------'
- anotherpaulg 2y agoLlama 3.1 405B instruct is #7 on aider's leaderboard, well behind Claude 3.5 Sonnet & GPT-4o. When using SEARCH/REPLACE to efficiently edit code, it drops to #11. https://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/ 77.4% claude-3.5-sonnet 75.2% DeepSeek Coder V2 (whole) 72.9% gpt-4o 69.9% DeepSeek Chat V2 0628 68.4% claude-3-opus-20240229 67.7% gpt-4-0613 66.2% llama-3.1-405b-instruct (whole)
- j_maffe 2y agoOrdinal value doesn't really matter in this case, especially when it's a categorically different option, access-wise. A 10% difference isn't bad at all.
- Jiahang 2y agoit is nice to see the 405b model is actually competitive against closed source frontier models But i just have M2pro may can't play it
- jiriro 2y agoCan this Llama process ~1GB of custom XML data? And answer queries like: Give all <myObject> which refer to <location> which refer to an Indo-European <language>.
- hrpnk 2y agoThe model's context is 128k tokens, so you'd have to split the data and analyze in chunks.