7 ms·
General overview below, as the pages don't seem to be working well Llama 4 Models: - Both Llama 4 Scout and Llama 4 Maverick use a Mixture-of-Experts (MoE)
by laborcontract 2y ago
General overview below, as the pages don't seem to be working well
Llama 4 Models:
- Both Llama 4 Scout and Llama 4 Maverick use a Mixture-of-Experts (MoE) design with 17B active parameters each.
- They are natively multimodal: text + image input, text-only output.
- Key achievements include industry-leading context lengths, strong coding/reasoning performance, and improved multilingual capabilities.
- Knowledge cutoff: August 2024.
Llama 4 Scout:
- 17B active parameters, 16 experts, 109B total.
- Fits on a single H100 GPU (INT4-quantized).
- 10M token context window
- Outperforms previous Llama releases on multimodal tasks while being more resource-friendly.
- Employs iRoPE architecture for efficient long-context attention.
- Tested with up to 8 images per prompt.
Llama 4 Maverick:
- 17B active parameters, 128 experts, 400B total.
- 1M token context window.
- Not single-GPU; runs on one H100 DGX host or can be distributed for greater efficiency.
- Outperforms GPT-4o and Gemini 2.0 Flash on coding, reasoning, and multilingual tests at a competitive cost.
- Maintains strong image understanding and grounded reasoning ability.
Llama 4 Behemoth (Preview):
- 288B active parameters, 16 experts, nearly 2T total.
- Still in training; not yet released.
- Exceeds GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on STEM benchmarks (e.g., MATH-500, GPQA Diamond).
- Serves as the “teacher” model for Scout and Maverick via co-distillation.
Misc:
- MoE Architecture: Only 17B parameters activated per token, reducing inference cost.
- Native Multimodality: Unified text + vision encoder, pre-trained on large-scale unlabeled data.
- qwertox 2y agoLlama 4 Scout, Maximum context length: 10M tokens. This is a nice development.
- lostmsu 2y agoHow did they achieve such a long window and what are the memory requirements to utilize it?
- deleted 2y ago[deleted]
- miven 2y agoAccording to [0] it's partly due to a key change they introduced in interleaving layers that use standard RoPE positional encodings and layers using what's called NoPE [1], not encoding positions at all and letting the model to figure those out on its own (this exclusively works because the LLMs are autoregressive, so the model can recognize an input token as being the very first by there not yet being any other tokens to attend to, and recursively deriving the position of the subsequent ones from that base case) [0] https://ai.meta.com/blog/llama-4-multimodal-intelligence/ https://ai.meta.com/blog/llama-4-multimodal-intelligence/ [1] https://arxiv.org/abs/2305.19466 https://arxiv.org/abs/2305.19466
- deleted 2y ago[deleted]
- lelandbatey 2y agoIs the recall and reasoning equally good across the entirety of the 10M token window? Cause from what I've seen many of those window claims equate to more like a functional 1/10th or less context length.
- Baeocystin 2y agoI assume they're getting these massive windows via RAG trickery, vectorization, and other tricks behind the curtain, became I've noticed the same as you- things start dipping in quality pretty quickly. Does anyone know if I am correct in my assumption?
- jimmyl02 2y agothe large context windows generally involve RoPE[0] which is a trick that allows the training window to be smaller but expand larger during inference. it seems like they have a new "iRoPE" which might have better performance? [0]https://arxiv.org/pdf/2104.09864 https://arxiv.org/pdf/2104.09864
- reissbaker 2y agoThere's no "RAG trickery" or vector search. They changed the way they encode positions such that in theory they're less sensitive to where the token appears in the string. That's similar to how previous long-context models worked as well, although the earlier iterations didn't work particularly well, as most have noticed; technically the model "worked" with longer contexts, but it would definitely get dumber. Still too early to tell how this newer variant works, although I'd assume it's at least somewhat better.
- jimmyl02 2y agothe needle in a haystack benchmark looks good but at this point I think we need new benchmarks to test actual understanding of content in such a large window.
- vessenes 2y agoIt’s going to take a while to see how good this window is for real use; they’ve used a couple new ideas to get to 10M token context. Right now the only really good long token model out there is Gemini Pro - and its effectiveness does start dropping maybe in the 200k token range. I imagine insiders at GOOG have access to more than the published 1M token range there. It will be fun to see what we get here, but I have no doubt the extra tokens will be useful - lots of use cases can do almost as well with summary-level accuracy memory.
- aimanbenbaha 2y agoI don't think RAG will survive this time
- drusepth 2y agoRAG still has lots of benefits for anyone paying per input token (e.g. over APIs).
- azinman2 2y agoNot to mention latency
- disgruntledphd2 2y agoAnd grounding for the model. Smaller models with tend to hallucinate a little less (anecdotally).
- inertiatic 2y ago4.8b words on English Wikipedia. Knowledge cutoff of 6 months. A valid use case is to search across Wikipedia and ground your answers. Trivially proves that RAG is still needed.
- acchow 2y agoThis is only for the small model. The medium model is still at 1M (like Gemini 2.5) Even if we could get the mid models to 10M, that's still a medium-sized repo at best. Repos size growth will also accelerate as LLMs generate more code. There's no way to catch up.
- gesman 2y agoRAG gets bigger as everyone else gets bigger. Flooding prompts with garbage is not a sound strategy...
- accrual 2y agoThanks for sharing this here. At first I loved the simple Apache-style directory listing, very classic and utilitarian way to navigate new information. Then I tried clicking the FAQ and it wouldn't load anything until I allowed two different sources of JavaScript.
- clueless 2y ago> Knowledge cutoff: August 2024. Could this mean training time is generally around 6 month, with 2 month of Q/A?
- bertil 2y agoCouldn’t you gradually include more recent documents as you train?
- soulofmischief 2y agoThat makes it harder to analyze the results of training and draw conclusions for the next round.
- changoplatanero 2y agoYou can do that but the amount of incremental data will be negligible compared to the rest of the data. Think of the knowledge cutoff more like a soft value.
- nickysielicki 2y agoIt scales depending on the dataset you want exposure on and the compute you have available, so any specific time box is kind of meaningless if you don’t know the rest of the inputs that went into it. The llama 3 paper went into a lot of this and how these decisions were made (see section 3 and onward): https://ai.meta.com/research/publications/the-llama-3-herd-of-models/ https://ai.meta.com/research/publications/the-llama-3-herd-o... tl;dr: llama 3 was 54 days, but it’s more complicated than that.
- jhugg 2y agoI wish my knowledge cutoff was August 2024.
- steenandersson 2y agoThis made me LOL louder than I have for a long time! Agree.
- InvOfSmallC 2y agoFor a super ignorant person: Both Llama 4 Scout and Llama 4 Maverick use a Mixture-of-Experts (MoE) design with 17B active parameters each Those experts are LLM trained on specific tasks or what?
- vessenes 2y agoThis was an idea that sounded somewhat silly until it was shown it worked. The idea is that you encourage through training a bunch of “experts” to diversify and “get good” at different things. These experts are say 1/10 to 1/100 of your model size if it were a dense model. So you pack them all up into one model, and you add a layer or a few layers that have the job of picking which small expert model is best for your given token input, route it to that small expert, and voila — you’ve turned a full run through the dense parameters into a quick run through a router and then a 1/10 as long run through a little model. How do you get a “picker” that’s good? Well, it’s differentiable, and all we have in ML is a hammer — so, just do gradient descent on the decider while training the experts! This generally works well, although there are lots and lots of caveats. But it is (mostly) a free lunch, or at least a discounted lunch. I haven’t seen a ton of analysis on what different experts end up doing, but I believe it’s widely agreed that they tend to specialize. Those specializations (especially if you have a small number of experts) may be pretty esoteric / dense in their own right. Anthropic’s interpretability team would be the ones to give a really high quality look, but I don’t think any of Anthropic’s current models are MoE. Anecdotally, I feel MoE models sometimes exhibit slightly less “deep” thinking, but I might just be biased towards more weights. And they are undeniably faster and better per second of clock time, GPU time, memory or bandwidth usage — on all of these - than dense models with similar training regimes.
- randomcatuser 2y agoyes, and it's on a per-layer basis, I think! So if the model has 16 transformer layers to go through on a forward pass, and each layer, it gets to pick between 16 different choices, that's like 16^16 possible expert combinations!
- Buttons840 2y ago
- kristopolous 2y ago17B puts it beyond the reach of a 4090 ... anybody do 4 bit quant on it yet?
- taneq 2y agoUnless something’s changed you will need the whole model on the HPU anyway, no? So way beyond a 4090 regardless.
- kristopolous 2y agoA habana just for inference? Are you sure? Also I see the 4 bit quants put it at a h100 which is fine ... I've got those at work. Maybe there will be distilled for running at home
- littlestymaar 2y agoYou can still offload most of the model to RAM and use the GPU for compute, but it's obviously much slower than what it would be if everything was on the GPU memory. see ktransformers: https://www.reddit.com/r/LocalLLaMA/comments/1jpi0n9/ktransformers_now_supports_multiconcurrency_and/ https://www.reddit.com/r/LocalLLaMA/comments/1jpi0n9/ktransf...
- kristopolous 2y agoI'm certainly not the brightest person in this thread but has there been effort to maybe bucket the computational cost of the model so that more expensive parts are on the gpu and less expensive parts are on the cpu?
- phonon 2y agoTake a look at https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/DeepseekR1_V3_tutorial.md https://github.com/kvcache-ai/ktransformers/blob/main/doc/en...
- reissbaker 2y ago
- ramshanker 2y agoI have a gut feeling, next in line will be 2 or more level of MoE. Further reducing the memory bandwidth and compute requirements. So top level MoE router decides which sub MoE to route.
- jamesblonde 2y agoThe solution to all problems in computer science is add a new level of indirection (or abstraction).
- brookst 2y agoExcept when the solution is to collapse abstraction in the name of efficiency.
- fsndz 2y agoNice release. I see that everyone is playing the differentiation game now: https://medium.com/thoughts-on-machine-learning/llama-4-and-the-differentiation-game-e21aeae59b7c https://medium.com/thoughts-on-machine-learning/llama-4-and-...
- MR4D 2y agoIf their knowledge cutoff is 8 months ago, then how on earth does Grok know things that happened yesterday? I would really love to know that.