24 ms·
Mistral AI Launches New 8x22B MOE Model
- swalsh 2y agoIs this Mistral large?
- varunvummadi 2y agoNot sure trying to download the torrent and checking it out
- fbdab103 2y agoFor those of us without twitter, how many GB is the model?
- KTibow 2y ago(hope this isn't against rules but) If you don't have Twitter, the magnet link is magnet:?xt=urn:btih:9238b09245d0d8cd915be09927769d5f7584c1c9&dn=mixtral-8x22b&tr=udp%3A%2F%http://2Fopen.demonii.com%3A1337%2Fannounce&tr=http%3A%2F%http://2Ftracker.opentrackr.org%3A1337%2Fannounce
- gjs278 2y ago[dead]
- confused_boner 2y ago262 gb
- fbdab103 2y agoOoof. I really need to pick up another HD, these model sizes are killer. Lacking a godly GPU, I will probably hold off for a quanitized version which has the potential to run okish on CPU or my modest GPU, but really appreciate the info.
- Jackson__ 2y agoUnlikely, this model has a max sequence length of 65k, while mistral large is 32k.
- varunvummadi 2y agoThey Just announced their new model on Twitter, which you can download using torrent
- deleted 2y ago[deleted]
- mlsu 2y ago8x22b. If this is as good as Mixtral 8x7b we are in for a wonderful time.
- cchance 2y agoI've heard command-r is first opensource to beat gpt4 in benchmarks
- varunvummadi 2y agoIt beats the old GPT4 version in lmsys benchmark you can check it out here https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboard https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar... but Command R is commercially licensed We can assume that mistral will do a better job.
- skissane 2y ago> but Command R is commercially licensed It is licensed under CC-BY-NC-4.0. That license means you are free to use, modify and redistribute it, so long as you aren't doing so "commercially". What exactly counts as "commercial" use is a complex legal question, and the answer may vary from jurisdiction to jurisdiction (different courts may interpret the phrase differently). But, for example, if you are just using it at home for private experimentation on your own personal time, with no plans to make money from doing so (whether now or in the future), I think pretty much everyone will agree that counts as "non-commercial". Other cases – e.g., if a government agency uses the software to provide some government function, is that "non-commercial"? – are far less clear. Those are really the kind of questions you need to ask a lawyer (which I am not).
- ryao 2y agoI am not a lawyer, but lately, I have been wondering whether the contra proferentem rule interacts with these licenses.
- 2y ago
- nen-nomad 2y agoMixtral 8x7b has been good to work with, and I am looking forward to trying this one as well.
- SushiHippie 2y ago[dupe] https://news.ycombinator.com/item?id=39986047 https://news.ycombinator.com/item?id=39986047 Which has the link to the tweet instead of the profile: https://twitter.com/MistralAI/status/1777869263778291896 https://twitter.com/MistralAI/status/1777869263778291896
- zmmmmm 2y agoA pre-Llama3 race for everyone to get their best small models on the table?
- swyx 2y agothis is likely v true given llama 3 rumored to release in next 2 weeks
- moffkalast 2y ago262 GB is not exactly small. But yes it seems they're all getting them out the door in case they end up being worse than llama-3 in which case it'll be too embarrassing to release later.
- hmottestad 2y agoSince it’s a MOE model it will only need to load a few of the 8 sub models into vram in order to answer a query. So it may look large, but I think a quantized model will easily fit on a Mac with 64GB of memory and maybe even a bit fewer bits and it’ll fit into 32GB. I think it might be the end for 24GB 4090 cards though :(
- brandall10 2y agoUnless something has changed, it needs to load the full 8 models at the same time. During inference it performs like a 2 x base model. Mixtral 7B @ 5 bit takes up over 30gb on my M3 Max. That's over 90 for this at the same quantization. Realistically you probably need a 128gb machine to run this with good results.
- fzzzy 2y agoA 4 bit quant of the new one would still be about 70 gb, so yeah. Gonna need a lot more ram.
- mark_l_watson 2y agoI think you are an optimist here. I can barely run mixtral-8x-7B on my M2 Pro 32G Mac, but I am grateful to be able to run it at all.
- tjtang2019 2y agoWhat are the advantages compared to GPT? Looking forward to using it!
- qball 2y ago>What are the advantages compared to GPT? It actually does what you tell it, and won't try to silently change your prompt to conform to a specific flavor of Californian hysterics, which is what OpenAI's products do. Also, since it's a local model, your queries aren't being datamined nor can access to the service be revoked on a whim.
- abdullahkhalids 2y agoWhy are some of their models open, and others closed? What is their strategy?
- kvmet 2y agoIt's gotta be either perceived value or training data/licensing restrictions.
- blackeyeblitzar 2y agoI am not sure why some are open and some are closed - if I had to speculate, it’s perhaps that the commercial models help fund the team. They come with safety features built-in as well as API-based access (instead of needing to self-host). They word their mission (https://mistral.ai/company/#missions https://mistral.ai/company/#missions) as follows: > Our mission is to make frontier AI ubiquitous, and to provide tailor-made AI to all the builders. This requires fierce independence, strong commitment to open, portable and customisable solutions, and an extreme focus on shipping the most advanced technology in limited time.
- unraveller 2y agoMistral have stated they want to chase the fine-tune dollar to support le research. We should get thrown a bone of hard to tune mid-range stuff occasionally. Especially when big announcements about small models are expected later in the week (llama3) or when haiku is stealing the thunder from mixtral 8x7b.
- Jackson__ 2y agoMy personal speculation is that their closed models are based on other companies' models. For example on EQbench[0], Miqu[1], a leaked continued pretrain based on LLama2, performs extremely similar to the mistral medium model their API offers. Maybe they're thinking it'd be bad PR for them to release models they didn't create from scratch, or there is some contractual obligation preventing the release. [0]https://eqbench.com/index.html https://eqbench.com/index.html [1]https://huggingface.co/miqudev/miqu-1-70b https://huggingface.co/miqudev/miqu-1-70b
- 2y ago
- angilly 2y agoThe lack of a corresponding announcement on their blog makes me worry about a Twitter account compromise and a malicious model. Any way to verify it’s really from them?
- swyx 2y agoyou must be new to mistral releases. they invented the magnet first blog later meta
- angilly 2y agoAt 3:30a France local? Alrighty. I still wait a lil bit ;)
- moralestapia 2y agoWhat could a malicious model do, though? Curse at you?
- Teever 2y agohttps://arstechnica.com/security/2024/03/hugging-face-the-github-of-ai-hosted-code-that-backdoored-user-devices/ https://arstechnica.com/security/2024/03/hugging-face-the-gi...
- Tiberium 2y agoNot .safetensors though
- deleted 2y ago[deleted]
- Aissen 2y agoExploit a memory safety issue in the tokenizer/or other parts of your LLM infra written in a native language.
- talsperre 2y agoRight on time as LLama 3 is released.
- jimmySixDOF 2y agoAnd the same day Google Gemini Pro gets almost complete open long context multimodal access and OpenAI upgrade to GPT4-Turbo it was a big day in general for news drops that's for sure!
- freeqaz 2y agoWhat's the easiest way to run this assuming that you have the weights and the hardware? Even if it's offloading half of the model to RAM, what tool do you use to load this? Ollama? Llama.cpp? Or just import it with some Python library? Also, what's the best way to benchmark a model to compare it with others? Are there any tools to use off-the-shelf to do that?
- varunvummadi 2y agoThe easiest is to use vllm (https://github.com/vllm-project/vllm https://github.com/vllm-project/vllm) to run it on a Couple of A100's, and you can benchmark this using this library (https://github.com/EleutherAI/lm-evaluation-harness https://github.com/EleutherAI/lm-evaluation-harness)
- sheepscreek 2y agoIn that regard, it’s even easier to use one Apple Studio with sufficient RAM and llama.cpp or even PyTorch for inference.
- fbdab103 2y agoI think the llamafile[0] system works the best. Binary works on the command line or launches a mini webserver. Llamafile offers builds of Mixtral-8x7B-Instruct, so presumably they may package this one up as well (potentially a quantized format). You would have to confirm with someone deeper in the ecosystem, but I think you should be able to run this new model as is against a llamafile? [0] https://github.com/Mozilla-Ocho/llamafile https://github.com/Mozilla-Ocho/llamafile
- noman-land 2y ago+1 on llamafile. You can point it to a custom model.
- jart 2y agollamafile author here. I'm downloading Mixtral 8x22b right now. I can't say for certain it'll work until I try it, but let's keep our fingers crossed! If not, we'll be shipping a release as soon as possible that gets it working. My recent work optimizing CPU evaluation https://justine.lol/matmul/ https://justine.lol/matmul/ may have come at just the right time. Mixtral 8x7b always worked best at Q5_K_M and higher, which is 31GB. So unless you've got 4x GeForce RTX 4090's in your computer, CPU inference is going to be the best chance you've got at running 8x22b at top fidelity.
- ein0p 2y agoTo this day 8x7b Mixtral remains the best model you can run on a single 48GB GPU. This has the potential to become the best model you can run on two such GPUs, or on an MBP with maxed out RAM, when 4-bit quantized.
- noman-land 2y agoMy first thought was how much RAM? Will it work on 64GB M1?
- ein0p 2y agoNope. Just the weights would take 88GB at 4 bit. 128GB MBP ought to be able to run it. If I were to guess, a version for Apple MLX should be available within a few days, for those of us fortunate enough to own such a thing.
- Art9681 2y agoIt’s already available. I had it running yesterday morning in an M3 MAX 128GB. I get about 6tps. https://www.reddit.com/r/LocalLLaMA/s/MSsrqWHYga https://www.reddit.com/r/LocalLLaMA/s/MSsrqWHYga
- jwitthuhn 2y agoIt is ~260GB with presumably fp16 weights. Should fit into 64GB at 3-bit quantization (~49GB). Edit: To add to this, I've had good luck getting solid output out of mixtral 8x7b at 3-bit, so that isn't small enough to completely kill the model's quality.
- deoxykev 2y ago4 bit quants should require 85GB VRAM, so this will fit nicely on 4x 24G consumer GPUs, plus some leftover for KV cache optimization.
- hedgehog 2y agoI've found the 2 bit quant of Mixtral 8x7B is usable for some purposes with an 8GB GPU. I'm curious how this new model will work in similar cheap 8-16GB GPU configurations.
- cjbprime 2y agoWouldn't expect that to work at all.
- hedgehog 2y agoOllama (which wraps llama.cpp) supports splitting a model across devices so you get some acceleration even on models too big to fit entirely in GPU memory.
- reissbaker 2y ago16GB will be way too small unfortunately — this has over 3x the param count, so at best you're looking at a 24GB card with extreme 2bit quantization. Really though if you're just looking to run models personally and not finetune (which requires monstrous amounts of VRAM), Macs are the way to go for this kind of mega model: Macs have unified memory between the GPU and CPU, and you can buy them with a lot of RAM. It'll be cheaper than trying to buy enough GPU VRAM. A Mac Studio with 192GB unified RAM is under $6k — two A6000s will run you over $9k and still only give you 96GB VRAM (and God help you if you try to build the equivalent system out of 4090s or A100s/H100s). Or just rent the GPU time as needed from cloud providers like RunPod, although that may or may not be what you're looking for.
- dannyw 2y agoYou can QLoRA decent models on 24GB VRAM. There’s also optimised kernels like Unsloth that are really VRAM efficient and good for hobbyists.
- nazka 2y agoOut of topic but are we now back at the same performance than ChatGPT 4 at the time people said it worked like magic (meaning before the nerf to make it more politically correct but making his performance crash)?
- segmondy 2y agoWith open models, yes we are at the performance of at least the first release of ChatGPT 4.
- sp332 2y agoCould you recommend one or a few in particular?
- sanjiwatsuki 2y agoThe current best open weights model is probably Cohere Command-R+. The memory requirements on it are quite high, though.
- bevekspldnw 2y agoI really want to see some benchmarks with performance weighted by energy use. I think Mistral 7B performance to watt would be the leader by a huge margin. On many tasks I get equal performance on zero shot classification tasks on Mistral than in bigger models.
- hmottestad 2y agoI’ve been testing a lot of LLMs on my MacBook and I would say that all of them are far away from being as good as GPT-4, at any time. Many are as good as GPT-3 though. There are also a lot of models that are fine tuned for specific tasks. Language support is one big thing that is missing from open models. I’ve only found one model that can do anything useful with Norwegian, which has never been an issue GPT-4.
- 2y ago
- aurareturn 2y agoMight be a dumb question but does this mean this model has 176B params?
- hovering_nox 2y ago8x7 had 46B or so.
- idiliv 2y agoIn Mixtral 8x7B, the 8 means that the model uses Mixture-of-Experts (MoE) layers with 8 experts. The 7B means that if you were to remove 7 of the 8 experts in each layer, then you would end up with a 7B model (which would have exactly the same architecture as Mistral 7B). Therefore, a 1x7B model has 7B params. An 8x7B model has 1 * 7B + (8-1) * sz_expert params, where sz_expert is some constant value that the MoE layers increase by when adding one expert. In the case of Mixtral 8x7B the model size is 46.3GB, so, sz_expert ≈ 5.6B. If these assumptions port over to 8x22B, then 8x22B has, at 281GB, sz_expert ≈ 13.8B.
- deleted 2y ago[deleted]
- idiliv 2y agoOh, and to answer your actual question: Assuming that the model is released with 16 bits per parameter, then it as 281GB / 16 bit = 140.5 parameters.
- KTibow 2y agoI tried to check this for myself. I agreed for the first one, (46.3 - 7) / 7 = 5.61b. The second one doesn't match up, (281 - 22) / 7 = 37b or (140.5 - 22) / 7 = 16.92b. Am I doing something wrong?
- idiliv 2y agoJust tried this again and I also arrive at 16.92B. Not sure what I did wrong the first time, thanks for double-checking this!
- wkat4242 2y agoWeird, the last post I see at that link is from the 8th of December 2023 and it's not about this. Edit: Ah, it's the wrong link. https://news.ycombinator.com/item?id=39986047 https://news.ycombinator.com/item?id=39986047 Thanks SushiHippie!
- intellectronica 2y agoIt's weird that more than a day after the weights dropped, there still isn't a proper announcement from Mistral with a model card. Nor is it available on Mistral's own platform.
- tosh 2y agoat least they confirmed it is Apache 2.0 https://twitter.com/arthurmensch/status/1778308399144333411 https://twitter.com/arthurmensch/status/1778308399144333411
- ZeljkoS 2y agoHere is the unofficial benchmark: https://huggingface.co/mistral-community/Mixtral-8x22B-v0.1/discussions/4 https://huggingface.co/mistral-community/Mixtral-8x22B-v0.1/...
- bevekspldnw 2y agoWish it had GPT-4, that’s the one to beat still.
- GuB-42 2y agoIt is there, not for all the benchmarks, but for those where it is included, GPT-4 scores much higher. Not surprising since GPT-4 is still state-of-the-art and much bigger. Where Mistral has been particularly impressive is when you take the size of the model into account.
- mirekrusin 2y agoGPT-4 is instruct tuned model, of course it's going to score higher, apples and oranges.
- bevekspldnw 2y agoYeah and the instruct tunes provided by Mistral on other models are pretty great.
- stainablesteel 2y agohas anyone had success making an auto-gpt concept for mistral/llama models? i haven't found one
- dkasper 2y agoHas anyone had success making an auto-gpt with any models? Besides toy use cases
- danenania 2y agoI built one using GPT-4[1]. It's not perfect but is working quite well and is now being used by hundreds of users, apart from me, to work on real, non-toy tasks. For example, I used it to build most of a production-ready AWS infrastructure (and accompanying deploy script) with the AWS CDK. I want to add Mistral support soon, probably via together.ai or a similar service. 1 - https://github.com/plandex-ai/plandex https://github.com/plandex-ai/plandex
- resource_waste 2y agoWhat is the excitement around models that arent as good as llama? This is clearly an inferior model that they are willing to share for marketing purposes. If it was an improvement over llama, sure, but it seems like just an ad for bad AI.
- cma 2y agoIt beats llama on the benchmark posted below (though maybe leaked into training data). But also you can run it on cheaper split up hardware with less individual vram than the big llama.
- zone411 2y agoWhat makes it you think it's not as good as LLaMA? It's likely much better. There are multiple open-weight models that are better than LLaMA 2 out there already.
- jeppebemad 2y agoWe use their earlier Mixtral model because it outperforms llama for our use case. They do not release full models for marketing purposes, though it definitely grabs attention! You may need to revise your views..
- Me1000 2y agoMixtral 7x8b was way better than llama2 70b and used less RAM and compute at the same time. This model is way better than llama. In fact I would go as far as saying llama2 isn’t that good compared to some of the most recent models.
- zone411 2y agoVery important to note that this is a base model, not an instruct model. Instruct fine-tuned models are what's useful for chat.
- haolez 2y agoWhat's the feeling of playing with a powerful base model? Will it just complete the prompt text like a continuation of it?
- MPSimmons 2y agoGenerally, yes, it literally just tries to predict the next token again and again and again. This model is apparently surprisingly good at chat, even though it is a base model, and will take part it it to some extent. It should be really interesting once it's fine-tuned.
- T3RMINATED 2y ago[dead]