26 ms·
Mistral "Mixtral" 8x7B 32k model [magnet]
- mareksotak 3y agoSome companies spend weeks on landing pages, demos and cute thought through promo videos and then there is Mistral, casually dropping a magnet link on Friday.
- tananaev 3y agoI'm sure it's also a marketing move to build a certain reputation. Looks like it's working.
- HlessClaudesman 3y agoNot geoblocking the entirety of Europe also makes them stand out like a ringmaster amongst clowns.
- moffkalast 3y agoWell they are French after all. They should be geoblocking the USA in response for a bit to make a point lol.
- fredoliveira 3y agoNot with their cap table, they won't ;-)
- peanuty1 3y agoGoogle Bard is still not available in Canada.
- oh_sigh 3y agoAre there some regulatory reasons why it would not be available? It seems weird if Google would intentionally block users merely to block them.
- mrandish 3y agoI think there are still some pretty onerous laws about French localization of products and services made available in the French-speaking part of Canada. Could be that...
- simonerlic 3y agoI originally thought so too, but as far as I know Bard is available in France- so I have a feeling that language isn't the roadblock here.
- dpwm 3y agoCan confirm Bard is available in France and the UI has been translated to French.
- wadefletch 3y agoThere's a proposed framework[1] in the EU that's rather restrictive. Seems like they're just not even bothering, perhaps to make a point. [1] https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai https://digital-strategy.ec.europa.eu/en/policies/regulatory...
- ComputerGuru 3y agoGoogle and Facebook were, up until just a couple of days ago, in a cold war within the Canadian government.
- peanuty1 3y agoGoogle made a deal to pay 100M/year to news organizations in Canada but Meta is continuing to block news links.
- noonething 3y agoneither is facebooks image one
- throwaway4aday 3y agotechnically, it is marketing but at this level marketing is indistinguishable from shipping
- tarruda 3y agoI'm curious about their business model.
- jorge-d 3y agoWell so far their business model seems to be mostly centered about raising money[1]. I do hope they succeed in becoming a succesful contender against OpenAI. [1] https://www.bloomberg.com/news/articles/2023-12-04/openai-rival-mistral-nears-2-billion-valuation-with-nvidia-funding https://www.bloomberg.com/news/articles/2023-12-04/openai-ri...
- udev4096 3y agohttps://archive.ph/4F3dT https://archive.ph/4F3dT
- nuz 3y agoThey can make plenty by offering consulting fees for finetuning and general support around their models.
- udev4096 3y agohttps://nitter.rawbit.ninja/MistralAI/status/1733150512395038967 https://nitter.rawbit.ninja/MistralAI/status/173315051239503...
- politician 3y agoHonest question: Why isn't this on Huggingface? Is this one a leaked model with a questionable training or alignment methodology? EDIT: I mean, I guess they didn't hack their own twitter account, but still.
- kcorbitt 3y agoIt'll be on Huggingface soon. This is how they dropped their original 7B model as well. It's a marketing thing, but it works!
- politician 3y ago@kcorbitt Low priority, probably not worth an email: Does using OpenPipe.ai to fine-tune a model result in a downloadable LoRA adapter? It's not clear from the website if the fine-tune is exportable.
- politician 3y agoAh, well, ok. I appreciate the torrent link -- much faster distribution.
- ComputerGuru 3y agoAlso more reliable. I had to write my own script to clone hf repos on Windows because git+lfs to an smb share would only partially download.
- tarruda 3y agoStill 7B, but now with 32k context. Looking forward to see how it compares with the previous one, and what the community does with it.
- MacsHeadroom 3y agoNot 7B, 8x7B. It will run with the speed of a 7B model while being much smarter but requiring ~24GB of RAM instead of ~4GB (in 4bit).
- dragonwriter 3y agoGiven the config parametes posted, its 2 experts per token, so the conputation cost per token should be the cost of the conponent that selects experts + 2× cost of a 7B model.
- stavros 3y agoYes, but I also care about "can I load this onto my home GPU?" where, if I need all experts for this to run, the answer is "no".
- MacsHeadroom 3y agoThe answer is yes if you have a 24GB GPU. Just wait for 4bit quantization. Or watch Tim Dettmers, who is releasing code to run Mixtral 8x7b in just 4GB of RAM.
- MacsHeadroom 3y agoAh good catch. Upon even closer examination, the attention layer (~2B params) is shared across experts. So in theory you would need 2B for the attention head + 5B for each of two experts in RAM. That's a total of 12B, meaning this should be able to be run on the same hardware as 13B models with some loading time between generations.
- deleted 3y ago[deleted]
- nulld3v 3y agoLooks to be Mixture of Experts, here is the params.json: { "dim": 4096, "n_layers": 32, "head_dim": 128, "hidden_dim": 14336, "n_heads": 32, "n_kv_heads": 8, "norm_eps": 1e-05, "vocab_size": 32000, "moe": { "num_experts_per_tok": 2, "num_experts": 8 } }
- sp332 3y agoI don't see any code in there. What runtime could load these weights?
- brucethemoose2 3y agoIts presumably llama just like Mistral. Everything open source is llama now. Facebook all but standardized the architecture. I dunno about the moe. Is there existing transformers code for that part? It kinda looks like there is based on the config.
- jasonjmcghee 3y agoMistral is not llama architecture. https://github.com/mistralai/mistral-src https://github.com/mistralai/mistral-src
- brucethemoose2 3y agoIts basically llama architecture, all but drop in compatible with llama runtimes.
- refulgentis 3y agoBecause it's JSON? :)
- sockaddr 3y agoWhat does expert mean in this context?
- YetAnotherNick 3y ago86 GB. So it's likely a Mixture of experts model with 8 experts. Exciting.
- tarruda 3y agoDamn, I was hoping it was still a single 7B model that I would be able to run on my GPU
- renonce 3y agoYou can, wait for a 4-bit quantized version
- tarruda 3y agoI only have a RTX 3070 with 8GB VRam. It can run quantized 7B models well, but this is 8 x 7B. Maybe an RTX 3090 with 24GB VRAM can do it.
- brucethemoose2 3y agoIt would be very tight. 8x7B 24GB (currently) has more overhead than 70B. Its theoretically doable, with quantization from the recent 2 bit quant paper and a custom implementation (in exllamav2?) EDIT: Actually the download is much smaller than 8x7B. Not sure how, but its sized more like a 30B, perfect for a 3090. Very interesting.
- burke 3y agoNapkin math: 7x(4/8)x8 is 28GB, and q4 uses a little more than just 4 bits per param, and there’s extra overhead for context, and the FFN to select experts is probably more on top of that. It would probably fit in 32GB at 4-bit but probably won’t run with sensible quantization/perf on a 3090/4090 without other tricks like offloading. Depending on how likely the same experts are to be chosen for multiple sequential tokens, offloading experts may be viable.
- espadrine 3y ago
- _uqgj 3y agomultimodal? 32k context is pretty impressive, curious to test instructability
- brucethemoose2 3y agoMistralLite is already 32K, and Yi 200K actually works pretty well out to at least 75K (the most I tested)
- civilitty 3y agoWhat kind of tests did you run out to that length? (Needle in haystack, summarization, structured data extraction, etc) What is the max number of tokens in the output?
- brucethemoose2 3y agoLong stories mostly, either novel or chat format. Sometimes summarization or insights, notably tests that you could't possible do with RAG chunking. Mostly short responses, not rewriting documents or huge code blocks or anything like that. MistralLite is basically overfit to summarize and retrieve in its 32K context, but its extremely good at that for a 7B. Its kinda useless for anything else. Yi 200K is... smart with the long context. An example I often cite is a Captain character in a story I 'wrote' with the llm. A Yi 200K finetune generated a debriefing for like 40K of context in a story, correctly assessing what plot points should be kept secret and making some very interesting deductions. You could never possibly do that with RAG on a 4K model, or even models that "cheat" with their huge attention like Anthropic. I test at 75K just because that's the most my 3090 will hold.
- kcorbitt 3y agoNo public statement from Mistral yet. What we know: - Mixture of Experts architecture. - 8x 7B parameters experts (potentially trained starting with their base 7B model?). - 96GB of weights. You won't be able to run this on your home GPU.
- tarruda 3y agoTheoretically it could fit into a single 24GB GPU if 4-bit quantized. Exllama v2 has even more efficient quantization algorithm, and was able to fit 70B models in 24GB gpu, but only with 2048 tokens of context.
- deleted 3y ago[deleted]
- coder543 3y ago> 96GB of weights. You won't be able to run this on your home GPU. This seems like a non-sequitur. Doesn't MoE select an expert for each token? Presumably, the same expert would frequently be selected for a number of tokens in a row. At that point, you're only running a 7B model, which will easily fit on a GPU. It will be slower when "swapping" experts if you can't fit them all into VRAM at the same time, but it shouldn't be catastrophic for performance in the way that being unable to fit all layers of an LLM is. It's also easy to imagine caching the N most recent experts in VRAM, where N is the largest number that still fits into your VRAM.
- tarruda 3y agoI will be super happy if this is true. Even if you can't fit all of them in the VRAM, you could load everything in tmpfs, which at least removes disk I/O penalty.
- cjbprime 3y agoJust mentioning in case it helps anyone out: Linux already has a disk buffer cache. If you have available RAM, it will hold on to pages that have been read from disk until there is enough memory pressure to remove them (and then it will only remove some of them, not all of them). If you don't have available RAM, then the tmpfs wouldn't work. The tmpfs is helpful if you know better than the paging subsystem about how much you really want this data to always stay in RAM no matter what, but that is also much less flexible, because sometimes you need to burst in RAM usage.
- deleted 3y ago[deleted]
- cloudhan 3y agoMight be the training code related with the model https://github.com/mistralai/megablocks-public/tree/pstock/mixtral https://github.com/mistralai/megablocks-public/tree/pstock/m...
- cloudhan 3y agoMixtral-8x7B support --> Support new model https://github.com/stanford-futuredata/megablocks/pull/45 https://github.com/stanford-futuredata/megablocks/pull/45
- sergiotapia 3y agoStuck on "Retrieving data" from the Magnet link and "Downloading metadata" when adding the magnet to the download list. I had to manually add these trackers and now it works: https://gist.github.com/mcandre/eab4166938ed4205bef4 https://gist.github.com/mcandre/eab4166938ed4205bef4
- sigmar 3y agoNot exactly similar companies in terms of their goals, but pretty hilarious to contrast this model announcement with Google's Gemini announcement two days ago.
- aubanel 3y agoMistral sure does not bother too much with explanations, but this style gives me much more confidence in the product than Google's polished, corporate, soulless announcement of Gemini!
- brucethemoose2 3y agoI will take weights over docs. Its does remind me how some Google employee was bragging that they disclosed the weights for the Gemini, and only the small mobile Gemini, as if that's a generous step over other companies.
- refulgentis 3y agoI don't think that's true, because quite simply, they have not. I am 100% in agreement with your viewpoint, but feel squeamish seeing an un-needed lie coupled to it to justify it. Just so much Othering these days.
- brucethemoose2 3y agoI was referencing this tweet: https://twitter.com/zacharynado/status/1732425598465900708 https://twitter.com/zacharynado/status/1732425598465900708 (Alt: https://nitter.net/zacharynado/status/1732425598465900708 https://nitter.net/zacharynado/status/1732425598465900708) That is fair though, this was an impulsive addition on my part.
- whimsicalism 3y agothey did not disclose the weights for any gemini, you must have misunderstood
- gitfan86 3y ago[flagged]
- danielbln 3y ago[flagged]
- udev4096 3y agobased mistral casually dropping a magnet link
- manojlds 3y agoGoogle - Fake demo Mistral - magnet link and that's it
- maremmano 3y agoDo you need some fancy announcement? let's do it the 90s way: https://twitter.com/erhartford/status/1733159666417545641/photo/1 https://twitter.com/erhartford/status/1733159666417545641/ph...
- deleted 3y ago[deleted]
- eurekin 3y agoI find that a way more bold and confident than dropping a obviously manipulated and unrealistic marketing page or video
- maremmano 3y agoFrankly I don't know why Google continues to act this way. Let's remind the "Google Duplex: A.I. Assistant Calls Local Businesses To Make Appointments" story. https://www.youtube.com/watch?v=D5VN56jQMWM https://www.youtube.com/watch?v=D5VN56jQMWM Not that this affects Google's user base in any way, at the moment.
- eurekin 3y agoThey obviously have both money and great talent. Maybe they put out minimal effort only for investors that expect their presence in consumer space?
- polygamous_bat 3y ago> Frankly I don't know why Google continues to act this way. Unfortunately, that's because they have Wall St. analysts looking at their videos who will (indirectly) determine how big of a bonus Sundar and co takes home at the end of the year. Mistral doesn't have to worry about that.
- eurekin 3y agoThis makes so much sense! Thanks
- brucethemoose2 3y agoIn other llm news, Mistral/Yi finetunes trained with a new (still undocumented) technique called "neural alignment" are blasting other models in the HF leaderboard. The 7B is "beating" most 70Bs. The 34B in testing seems... Very good: https://huggingface.co/fblgit/una-xaberius-34b-v1beta https://huggingface.co/fblgit/una-xaberius-34b-v1beta https://huggingface.co/fblgit/una-cybertron-7b-v2-bf16 https://huggingface.co/fblgit/una-cybertron-7b-v2-bf16 I mention this because it could theoretically be applied to Mistral Moe. If the uplift is the same as regular Mistral 7B, and Mistral Moe is good, the end result is a scary good model. This might be an inflection point where desktop-runnable OSS is really breathing down GPT-4's neck.
- deleted 3y ago[deleted]
- _boffin_ 3y agoInteresting. One thing i noticed is that Mistral has a `max_position_embeddings` of ~32k while these have it at 4096. Any thoughts on that?
- brucethemoose2 3y agoIs complicated. The 7B model (cybertron) is trained on Mistral. Mistral is technically a 32K model, but it uses a sliding window beyond 32K, and for all practical purposes in current implementations it behaves like an 8K model. The 34B model is based on Yi 34B, which is inexplicably marked as a 4K model in the config but actually works out to 32K if you literally just edit that line. Yi also has a 200K base model... and I have no idea why they didn't just train on that. You don't need to finetune at long context to preserve its long context ability.
- ComputerGuru 3y agoDid you mean "but it uses a sliding window beyond" *8K*? Because I don't understand how the sentence would work otherwise.
- MyFirstSass 3y agoHot take but Mistral 7B is the actual state of the art of LLM's. ChatGPT 4 is amazing yes and i've been a day 1 subscriber, but it's huge, runs on server farms far away and is more or less a black box. Mistral is tiny, and amazingly coherent and useful for it's size for both general questions and code, uncensored, and a leap i wouldn't have believed possible in just a year. I can run it on my Macbook Air at 12tkps, can't wait to try this on my desktop.
- tarruda 3y ago> I can run it on my Macbook Air at 12tkps, can't wait to try this on my desktop. That seems kinda low, are you using Metal GPU acceleration with llama.cpp? I don't have a macbook, but saw some of the llama.cpp benchmarks that suggest it can reach close to 30tk/s with GPU acceleration.
- MyFirstSass 3y agoThanks for the tip. I'm on the M2 Air with 16 GB's of ram. If anyone has faster than 12tkps on Air's let me know. I'm using the LM Studio GUI over llama.cpp with the "Apple Metal GPU" option. Increasing CPU threads seemingly does nothing either without metal. Ram usage hovers at 5.5GB with a q5_k_m of Mistral.
- M4v3R 3y agoTry different quantization variations. I got vastly different speeds depending on which quantization I chose. I believe q4_0 worked very well for me. Although for a 7B model q8_0 runs just fine too with better quality.
- ukuina 3y agoLlamaFile typically outperforms LM Studio and even Ollama.
- andy_xor_andrew 3y agoI am with you on this. Mistral 7B is amazingly good. There are finetunes of it (the Intel one, and Berkeley Starling) that feel like they are within throwing distance of gpt3.5T... at only 7B! I was really hoping for a 13B Mistral. I'm not sure if this MOE will run on my 3090 with 24GB. Fingers crossed that quantization + offloading + future tricks will make it runnable.
- seydor 3y agolooks like they're too busy being awesome. i need a fake video to understand this! What memory will this need? I guess it won't run on my 12GB of vram "moe": {"num_experts_per_tok": 2, "num_experts": 8} I bet many people will re-discover bittorrent tonight
- brucethemoose2 3y agoLooks like it will squeeze into 24GB once the llama runtimes work it out. Its also a good candidate for splitting across small GPUs, maybe. One architecture I can envision is hosting prompt ingestion and the "host" model on the GPU and the downstream expert model weights on the CPU /IGP. This is actually pretty efficient, as the CPU/IGP is really bad at the prompt ingestion but reasonably fast at ~14B token generation. Llama.cpp all but already does this, I'm sure MLC will implement it as well.
- syntaxing 3y agoBitTorrent was the craze when llama was leaked on torrent. Then Facebook started taking down all huggingface repos and a bunch of people transitioned to torrent released temporarily. llama 2 changed all this but it was a fun time.
- cuuupid 3y agoStark contrast with Google's "all demo no model" approach from earlier this week! Seems to be trained off Stanford's Megablocks: https://github.com/mistralai/megablocks-public https://github.com/mistralai/megablocks-public
- BryanLegend 3y agoAndrej Karpathy's take: New open weights LLM from @MistralAI params.json: - hidden_dim / dim = 14336/4096 => 3.5X MLP expand - n_heads / n_kv_heads = 32/8 => 4X multiquery - "moe" => mixture of experts 8X top 2 Likely related code: https://github.com/mistralai/megablocks-public https://github.com/mistralai/megablocks-public Oddly absent: an over-rehearsed professional release video talking about a revolution in AI. If people are wondering why there is so much AI activity right around now, it's because the biggest deep learning conference (NeurIPS) is next week. https://twitter.com/karpathy/status/1733181701361451130 https://twitter.com/karpathy/status/1733181701361451130
- henrysg 3y ago> Oddly absent: an over-rehearsed professional release video talking about a revolution in AI.
- crakenzak 3y ago> it's because the biggest deep learning conference (NeurIPS) is next week. Can we expect some big announcements (new architectures, models, etc) at the conference from different companies? Sorry, not too familiar what the culture for research conferences is.
- jbarrow 3y agoTypically not. Google as an example: the transformer paper (Vaswani et al., 2017) was arxiv'd in June of 2017, and NeurIPS (the conference in which it was published) was in December of that year; BERT (Devlin et al., 2019) was similarly arxiv'd before publication. Recent announcements from companies tend to be even more divorced from conference dates, as they release anemic "Technical Reports" that largely wouldn't pass muster in a peer review.
- GaggiX 3y ago>-hidden_dim / dim = 14336/4096 => 3.5X MLP expand >- n_heads / n_kv_heads = 32/8 => 4X These two are exactly the same as the old Mistral-7B
- 3y ago
- ahmetkca 3y agoLet’s go multimodal
- Jayakumark 3y agohttps://huggingface.co/someone13574/mixtral-8x7b-32kseqlen https://huggingface.co/someone13574/mixtral-8x7b-32kseqlen
- noahzhang 3y ago[dead]
- maremmano 3y agoWho know if I can run this on MBC Pro M3 max 128gb? at what TPS?
- deoxykev 3y agoI would like to know this as well.
- M4v3R 3y agoBig chance that you’ll be able to run it using Ollama app soon enough.
- marci 3y agoIf I understand correctly: RAM Wise, you can easily run a 70b with 128GB, 8x7B is obviously less than that. Compute wise, I suppose it would be a bit slower than running a 13b. edit: "actually", I think it might be faster than a 13b. 8 random 7b ~= 115GB, Mixtral is under 90. I will have to wait for more info/understanding.
- treprinum 3y agoI would say so based on LLaMA 2 70B; if it's 8x inference in MoE then I guess you'd see <20 tokens/sec?
- asolidtime1 3y agohttps://huggingface.co/someone13574/mixtral-8x7b-32kseqlen/blob/main/RELEASE https://huggingface.co/someone13574/mixtral-8x7b-32kseqlen/b... Holy shit, this is some clever marketing. Kinda wonder if any of their employees were part of the warez scene at some point.
- userbinator 3y agoThey certainly got that aesthetic right; the only thing that stands out (but might be a necessity) is using real names instead of handles.
- poulpy123 3y agois it eight 7b models in a trench coat ?
- fortunefox 3y agoReleasing a model with a magnet link and some ascii art gives me way more confidence in the product than any OpenAI blog post ever could. Excited to play with this once it's somewhat documented on how to get it running on a dual 4090 Setup.
- smlacy 3y agohttps://nitter.net/MistralAI/status/1733150512395038967 https://nitter.net/MistralAI/status/1733150512395038967
- leobg 3y agoI love Mistral. It’s crazy what can be done with this small model and 2 hours of fine tuning. Chatbot with function calling? Check. 90 +% accuracy multi label classifier, even when you only have 15 examples for each label? Check. Craaaazy powerful.
- leodriesch 3y agoCould you link me to a finetune optimized for function calling? I was looking for one a few weeks ago but did not find any.
- leobg 3y agoSee sibling comment.
- jeanloolz 3y agoCan you point me to a function calling fine tune mistral model? This is the only feature that keeps me from migrating away from openai. I searched a few time but could not find anything in HG
- leobg 3y agoCan’t share the model, since it was trained for a client. I don’t know if any public datasets exist. But Mistral will learn what you throw at it. So if you build a dataset of chat conversations that contains, say, answers in the form of {“answer”:”The answer”, “image”:”Prompt for stable diffusion”}, you’ll get a model that can generate images, and also will know when to use that capability. It’s insane how well that works.
- _fizz_buzz_ 3y agoDoes anybody have a tutorial or documentation how I can run this and play around with this locally. A „getting started“ guide of sorts?
- 0cf8612b2e1e 3y agoEven better if a llamafile gets released.
- deleted 3y ago[deleted]
- lxe 3y agoIf anyone can help running this, would be appreciated. Resources so far: - https://github.com/dzhulgakov/llama-mistral https://github.com/dzhulgakov/llama-mistral
- lagniappe 3y agoMagnet link says invalid for me
- stevebmark 3y agoMistral Mixtral Model Magnet Mistral Mixtral Model Magnet Mistral Mixtral Model Magnet
- jpdus 3y agoWe now have a (experimental) working HF version here: https://huggingface.co/DiscoResearch/mixtral-7b-8expert https://huggingface.co/DiscoResearch/mixtral-7b-8expert
- balnazzar 3y agoMight be relevant: https://twitter.com/dzhulgakov/status/1733217065811742863 https://twitter.com/dzhulgakov/status/1733217065811742863. Anyway, if the vanilla version requires 2x80gb cards, I wonder how would it run on a M2 Ultra 192gb Mac Studio. Anyone having the machine could try?
- yodsanklai 3y agoCan anyone explain what this means?
- ukuina 3y agoPossibly a huge leap forward in open-source model capability. GPT4's prowess supposedly comes from strong dataset + RLHF + MoE (Mixture of Experts). Mixtral brings MoE to an already-powerful model.
- dzhulgakov 3y agoYou can try Mixtral live at https://app.fireworks.ai/ https://app.fireworks.ai/ (soon to be faster too) Warning: the implementation might be off as there's no official one. We at Fireworks tried to reverse-engineer model architecture today with the help of awsome folks from the community. The generations look reasonably good, but there might be some details missing. If you want to follow the reverse-engineering story: https://twitter.com/dzhulgakov/status/1733330954348085439 https://twitter.com/dzhulgakov/status/1733330954348085439
- noahzhang 3y ago[dead]
- swah 3y agoKinda following all this stuff from outside w/o really understanding, but why are these things released like this, instead of "competing ChatGPTs apps" with higher and higher quality/costs? Could be open sourced but also hosted version that is maybe 5 usd/minute - if the results are great I guess people would pay the fair price... Is it mainly because its hard to apply the limitations so that it doesn't spit out bomb making instructions?