21 ms·
Mistral NeMo
- pantulis 2y agoDoes it have any relation to Nvidia's Nemo? Otherwise, it's unfortunate naming
- markab21 2y agoIt looks like it was built jointly with nvidia: https://huggingface.co/nvidia/Mistral-NeMo-12B-Instruct https://huggingface.co/nvidia/Mistral-NeMo-12B-Instruct
- refulgentis 2y agoClick the link, read the first sentence.
- pantulis 2y agoYeah, not my brightest HN moment, to be honest.
- SubiculumCode 2y agoAt least you didn't ask about finding a particular fish.
- deleted 2y ago[deleted]
- yjftsjthsd-h 2y ago> Today, we are excited to release Mistral NeMo, a 12B model built in collaboration with NVIDIA. Mistral NeMo offers a large context window of up to 128k tokens. Its reasoning, world knowledge, and coding accuracy are state-of-the-art in its size category. As it relies on standard architecture, Mistral NeMo is easy to use and a drop-in replacement in any system using Mistral 7B. > We have released pre-trained base and instruction-tuned checkpoints checkpoints under the Apache 2.0 license to promote adoption for researchers and enterprises. Mistral NeMo was trained with quantisation awareness, enabling FP8 inference without any performance loss. So that's... uniformly an improvement at just about everything, right? Large context, permissive license, should have good perf. The one thing I can't tell is how big 12B is going to be (read: how much VRAM/RAM is this thing going to need). Annoyingly and rather confusingly for a model under Apache 2.0, https://huggingface.co/mistralai/Mistral-Nemo-Instruct-2407 https://huggingface.co/mistralai/Mistral-Nemo-Instruct-2407 refuses to show me files unless I login and "You need to agree to share your contact information to access this model"... though if it's actually as good as it looks, I give it hours before it's reposted without that restriction, which Apache 2.0 allows.
- xena 2y agoEasy head math: parameter count times parameter size plus 20-40% for inference slop space. Anywhere from 8-40GB of vram required depending on quantization levels being used.
- imtringued 2y agoThey did quantization aware training for fp8 so you won't get any benefits from using more than 12GB of RAM for the parameters. What you might be using more RAM is the much bigger context window.
- exe34 2y agotensors look about 20gb. not sure what that's like in vram.
- kelsey98765431 2y agosame size
- renewiltord 2y agoAccording to nvidia https://blogs.nvidia.com/blog/mistral-nvidia-ai-model/ https://blogs.nvidia.com/blog/mistral-nvidia-ai-model/ it was made to fit on a 4090 so it should work with 24 GB.
- bernaferrari 2y agoif you want to be lazy, 7b = 7gb of vRAM, 12b = 12gb of vRAM, but quantizing you might be able to do with with ~6-8. So any 16gb Macbook could run it (but not much else).
- hislaziness 2y agoisn't it 2 bytes (fp16) per param. so 7b = 14 GB+some for inference?
- 2y ago
- Workaccount2 2y agoIs "Parameter Creep" going to becomes a thing? They hold up Llama-8b as a competitor despite NeMo having 50% more parameters. The same thing happened with gemma-27b, where they compared it to all the 7-9b models. It seems like an easy way to boost benchmarks while coming off as "small" at first glance.
- eyeswideopen 2y agoAs written here: https://huggingface.co/nvidia/Mistral-NeMo-12B-Instruct https://huggingface.co/nvidia/Mistral-NeMo-12B-Instruct "It significantly outperforms existing models smaller or similar in size." is a statement that goes in that direction and would allow the comparison of a 1.7T param model with a 7b one
- voiper1 2y agoOddly, they are only charging slightly more for their hosted version: open-mistral-7b is 25c/m tokens open-mistral-nemo-2407 is 30c/m tokens https://mistral.ai/technology/#pricing https://mistral.ai/technology/#pricing
- dannyw 2y agoPossibly a NVIDIA subsidy. You run NEMO models, you get cheaper GPUs.
- Palmik 2y agoThey specifically call out fp8 aware training and TensoRT LLM is really good (efficient) with fp8 inference on H100 and other hopper cards. It's possible that they run the 7b natively in fp16 as smaller models suffer more from even "modest" quantization like this.
- causal 2y agoYeah it will be interesting to see if we ever settle on standard sizes here. My preference would be: - 3B for CPU inference or running on edge devices. - 20-30B for maximizing single consumer GPU potential. - 70B+ for those who can afford it. 7-9B never felt like an ideal size.
- pants2 2y agoExciting, I think 12B is the sweet spot for running locally - large enough to be useful, fast enough to run on a decent laptop.
- _flux 2y agoHow much memory does employing the complete 128k window take, though? I've sadly noticed that it can take a significant amount of VRAM to use a larger context window. edit: e.g. I wouldn't know the correct parameters for this calculator, but going from 8k window to 128k window goes from 1.5 GB to 23 GB: https://huggingface.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator https://huggingface.co/spaces/NyxKrage/LLM-Model-VRAM-Calcul...
- azeirah 2y agoIn practice, it's fine to stick with "just" 8k or 16k or 32k. If you're working with data of over 128k tokens I'd personally not recommend using an open model anyway unless you know what you're doing. The models are kinda there, but the hardware mostly isn't. This is only realistic right now for people with those unified memory MacBook or for enthusiasts with Epyc servers or a very high end workstation built for inference. Anything above that I don't consider "consumer" inference
- mythz 2y agoIMO Google's Gemma2 27B [1] is the sweet spot for running locally on commodity 16GB VRAM cards. [1] https://ollama.com/library/gemma2:27b https://ollama.com/library/gemma2:27b
- minimaxir 2y ago> Mistral NeMo uses a new tokenizer, Tekken, based on Tiktoken, that was trained on over more than 100 languages, and compresses natural language text and source code more efficiently than the SentencePiece tokenizer used in previous Mistral models. Does anyone have a good answer why everyone went back to SentencePiece in the first place? Byte-pair encoding (which is what tiktoken uses: https://github.com/openai/tiktoken https://github.com/openai/tiktoken) was shown to be a more efficient encoding as far back as GPT-2 in 2019.
- rockinghigh 2y agoThe SentencePiece library also implements Byte-pair-encoding. That's what the LLaMA models use and the original Mistral models were essentially a copy of LLaMA2.
- zwaps 2y agoSentencePiece is not a different algorithm to WordPiece or BPE, despite its naming. One of the main pulls of the SentencePiece library was the pre-tokenization being less reliant on white space and therefore more adaptable to non Western languages.
- numeri 2y agoSentencePiece is a tool and library for training and using tokenizers, and supports two algorithms: Byte-Pair Encoding (BPE) and Unigram. You could almost say it is the library for tokenizers, as it has been standard in research for years now. Tiktoken is a library which only supports BPE. It has also become synonymous with the tokenizer used by GPT-3, ChatGPT and GPT-4, even though this is actually just a specific tokenizer included in tiktoken. What Mistral is saying here (in marketing speak) is that they trained a new BPE model on data that is more balanced multilingually than their previous BPE model. It so happens that they trained one with SentencePiece and the other with tiktoken, but that really shouldn't make any difference in tokenization quality or compression efficiency. The switch to tiktoken probably had more to do with latency, or something similar.
- p1esk 2y agoInteresting how it will compete with 4o mini.
- pixelatedindex 2y agoPardon me if this is a dumb question, but is it possible for me to download these models into my computer (I have a 1080ti and a [2|3]070ti) and generate some sort of api interface? That way I can write programs that calls this API, and I find this appealing. EDIT: This a 1W light bulb moment for me, thank you!
- bezbac 2y agoAFAIK, Ollama supports most of these models locally and will expose a REST API[0] [0]: https://github.com/ollama/ollama/blob/main/docs/api.md https://github.com/ollama/ollama/blob/main/docs/api.md
- kanwisher 2y agollama.cpp or ollama both have apis for most models
- codetrotter 2y agoI’d probably check https://ollama.com/library?q=Nemo https://ollama.com/library?q=Nemo in a couple of days. My guess is that by then ollama will have support for it. And you can then run the model locally on your machine with ollama.
- hedgehog 2y agoAdding to this: If the default is too slow look at the more heavily quantized versions of the model, they are smaller at moderate cost in output quality. Ollama can split models between GPU and host memory but the throughput dropoff tends to be pretty severe.
- andrethegiant 2y agoWhy would it take a couple days? Is it not a matter of uploading the model to their registry, or are there more steps involved than that?
- HanClinto 2y ago
- saberience 2y agoTwo questions: 1) Anyone have any idea of VRAM requirements? 2) When will this be available on ollama?
- causal 2y ago1) Rule of thumb is # of params = GB at Q8. So a 12B model generally takes up 12GB of VRAM at 8 bit precision. But 4bit precision is still pretty good, so 6GB VRAM is viable, not counting additional space for context. Usually about an extra 20% is needed, but 128K is a pretty huge context so more will be needed if you need the whole space.
- alecco 2y agoThe model has 12 billion parameters and uses FP8, so 1 byte each. With some working memory I'd bet you can run it on 24GB. > Designed to fit on the memory of a single NVIDIA L40S, NVIDIA GeForce RTX 4090 or NVIDIA RTX 4500 GPU
- jorgesborges 2y agoI’m AI stupid. Does anyone know if training on multiple languages provides “cross-over” — so training done in German can be utilized when answering a prompt in English? I once went through various Wikipedia articles in a couple languages and the differences were interesting. For some reason I thought they’d be almost verbatim (forgetting that’s not how Wikipedia works!) and while I can’t remember exactly I felt they were sometimes starkly different in tone and content.
- bernaferrari 2y agono, it is basically an 'auto-correct' spell checker from the phone. It only knows what it was trained on. But it has been shown that a coding LLM that has never seen a programming language or a library can "learn" a new one faster than, say, a generic LLM.
- StevenWaterman 2y agoThat's not true, LLMs can answer questions in one language even if they were only trained on that data in another language. IE you train an LLM on both English and French in general, but only teach it a specific fact in French, it can give you that fact in English
- hdhshdhshdjd 2y agoYou, you can write a prompt in English, give it French, and get an accurate answer in English even with the original Mistral. Still blows my mind we came so far so fast.
- miki123211 2y agoGenerally yes, with caveats. There was some research showing that training a model on facts like "the mother of John Smith is Alice" but in German allowed it to answer questions like "who's the mother of John Smith", but not questions like "what's the name of Alice's child", regardless of language. Not sure if this holds at larger model sizes though, it's the sort of problem that's usually fixable by throwing more parameters at it. Language models definitely do generalize to some extend and they're not "stochastic parrots" as previously thought, but there are some weird ways in which we expect them to generalize but they don't.
- madeofpalk 2y agoI find it interesting how coding/software development still appears to be the one category that these most popular models release specialised models for. Where's the finance or legal models from Mistral or Meta or OpenAI? Perhaps it's just confirmation bias, but programming really does seem to be the ideal usecase for LLMs in a way that other professions just haven't been able to crack. Compared to other types of work, it's relatively more straight forward to tell if code is "correct" or not.
- a2128 2y agoCoding models solve a clear problem and have a clear integration into a developer's workflow - it's like your own personal StackOverflow and it can autocomplete code for you. It's not as clear when it comes to finance or legal, you wouldn't want to rely on an AI that may hallucinate financial numbers or laws. These other professions are also a lot slower to react to change, compared to software development where people are already used to learning new frameworks every year
- MikeKusold 2y agoThose are regulated industries, where as software development is not. An AI spitting back bad code won't compile. An AI spitting back bad financial/legal advice bankrupts people.
- knicholes 2y agoGenerally I agree! I saw a guy shamefully admit he didn't read the output carefully enough when using generated code (that ran), but there was a min() instead of a max(), and it messed up a month of his metrics!
- troupo 2y agoThe explanation is easier, I think. Consider what data these models are trained on, and who are the immediate developers of these models. The models are trained on a vast set of whatever is available on the internet. They are developed by tech people/programmers who are surprisingly blind to their own biases and interests. There's no surprise that one of the main things they want to try and do is programming, using vast open quantities of Stack Overflow, GitHub and various programming forums. For finance and legal you need to: - think a bit outside the box - be interested in finance and legal - be prepared to carry actual legal liability for the output of your models
- simonw 2y agoI wonder why Mistral et al don't prepare GGUF versions of these for launch day? If I were them I'd want to be the default source of the versions of my models that people use, rather than farming that out to whichever third party races to publish the GGUF (and other formats) first.
- dannyw 2y agoI think it's actually reasonable to leave some opportunities to the community. It's an Apache 2.0 model. It's meant for everyone to build upon freely.
- a2128 2y agollama.cpp is still under development and they sometimes come out with breaking changes or new quantization methods, and it can be a lot of work to keep up with these changes as you publish more models over time. It's easier to just publish a standard float32 safetensors that works with PyTorch, and let the community deal with other runtimes and file formats. If it's a new architecture, then there's also additional work needed to add support in llama.cpp, which means more dev time, more testing, and potentially loss of surprise model release if the development work has to be done out in the open
- sroussey 2y agoSame could be said for onnx. Depends on which community you are in as to what you want.
- Patrick_Devine 2y ago
- mcemilg 2y agoI believe that if Mistral is serious about advancing in open source, they should consider sharing the corpus used for training their models, at least the base models pretraining data.
- wongarsu 2y agoI doubt they could. Their corpus almost certainly is mostly composed of copyrighted material they don't have a license for. It's an open question whether that's an issue for using it for model training, but it's obvious they wouldn't be allowed to distribute it as a corpus. That'd just be regular copyright infringement. Maybe they could share a list of the content of their corpus. But that wouldn't be too helpful and makes it much easier for all affected parties to sue them for using their content in model training.
- gooob 2y agono, not the actual content, just the titles of the content. like "book title" by "author". the tool just simply can't be taken seriously by anyone until they release that information. this is the case for all these models. it's ridiculous, almost insulting.
- candiddevmike 2y agoThey can't release it without admitting to copyright infringement.
- regularfry 2y agoThey can't do it without getting sued for copyright infringement. That's not quite the same.
- bilbo0s 2y agoUh.. That would almost be worse. All copyright holders would need to do is search a list of titles if I'm understanding your proposal correctly. The idea is not to get sued.
- alecco 2y agoNvidia has a blogpost about Mistral Nemo, too. https://blogs.nvidia.com/blog/mistral-nvidia-ai-model/ https://blogs.nvidia.com/blog/mistral-nvidia-ai-model/ > Mistral NeMo comes packaged as an NVIDIA NIM inference microservice, offering performance-optimized inference with NVIDIA TensorRT-LLM engines. > *Designed to fit on the memory of a single NVIDIA L40S, NVIDIA GeForce RTX 4090 or NVIDIA RTX 4500 GPU*, the Mistral NeMo NIM offers high efficiency, low compute cost, and enhanced security and privacy. > The model was trained using Megatron-LM, part of NVIDIA NeMo, with 3,072 H100 80GB Tensor Core GPUs on DGX Cloud, composed of NVIDIA AI architecture, including accelerated computing, network fabric and software to increase training efficiency.
- k__ 2y agoWhat's the reason for measuring the model size in context window length and not GB? Also, are these small models OSS? Easier self hosting seems to be the main benefo for small models.
- simion314 2y ago>What's the reason for measuring the model size in context window length and not GB? there are 2 different things. The context window is how many tokens ii's context can contain, so on a big model you could put in the context a few books and articles and then start your questions, on a small context model you can start a conversation and after a short time it will start forgetting eh first prompts. Big context will use more memory and will cost on performance but imagine you could give it your entire code project and then you can ask it questions, so often I know there is some functions already there that does soemthing but I can't remember the name.
- kaoD 2y agoI suspect you might be confusing the numbers: 12B (which is the very first number they give) is not context length, it's parameter count. The reason to use parameter count is because final size in GB depends on quantization. A 12B model at 8 bit parameter width would be 12Gbytes (plus some % overhead), while at 16 bit would be 24Gbytes. Context length here is 128k which is orthogonal to model size. You can notice the specify both parameters and context size because you need both to characterize an LLM. It's also interesting to know what parameter width it was trained on because you cannot get more information by "quantizing wider" -- it only makes sense to quantize into a narrower parameter width to save space.
- k__ 2y agoAh, yes. Thanks, I confused those numbers!
- yjftsjthsd-h 2y ago> Also, are these small models OSS? From the very first paragraph on the page: > released under the Apache 2.0 license.
- LoganDark 2y agoIs the base model unaligned? Disappointing to see alignment from allegedly "open" models.
- xena 2y agoThe reason that companies align models is so that they don't get on the front page of the new york times with a headline like "Techaro's AI model used by terrorists to build a pipe bomb that destroyed the New York Stock Exchange datacentre".
- bugglebeetle 2y agoInterested in the new base model for fine tuning. Despite Llama3 being a better instruct model overall, it’s been highly resistant to fine-tuning, either owing to some bugs or being trained on so much data (ongoing debate about this in the community). Mistral’s base model are still best in class for small model you can specialize.
- obblekk 2y agoWorth noting this model has 50% more parameters than llama3. There are performance gains but some of the gains might be from using more compute rather than performance per unit compute.
- dpflan 2y agoThese big models are getting pumped out like crazy, that is the business of these companies. But basically, it feels like private/industry just figured out how to scale up a scalable process (deep learning), and it required not $M research grants but $BB "research grants"/funding, and the scaling laws seem to be fun to play with and tweak more interesting things out of these and find cool "emergent" behavior as billions of data points get correlated. But pumping out models and putting artifacts on HuggingFace, is that a business? What are these models being used for? There is a new one at a decent clip.
- hdhshdhshdjd 2y agoI don’t see any indication this beats Llama3 70B, but still requires a beefy GPU, so I’m not sure the use case. I have an A6000 which I use for a lot of things, Mixtral was my go-to until Llama3, then I switched over. If you could run this on say, stock CPU that would increase the use cases dramatically, but if you still need a 4090 I’m either missing something or this is useless.
- azeirah 2y agoYou don't need a 4090 at all. 16 bit requires about 24GB of VRAM, 8bit quants (99% same performance) requires only 12GB of VRAM. That's without the context window, so depending on how much context you want to use you'll need some more GB. That is, assuming you'll be using llama.cpp (which is standard for consumer inference. Ollama is also llama.cpp, as is kobold) This thing will run fine on a 16GB card, and a q6 quantization will run fine on a 12GB card. You'll still get good performance on an 8GB card with offloading, since you'll be running most of it on the gpu anyway.
- reissbaker 2y agoComparing this to 70b doesn't make sense: this is a 12b model, which should easily fit on consumer GPUs. A 70b will have to be quantized to near-braindead to fit on a consumer GPU; 4bit is about as small as you can go without serious degradation, and 70b quantized to 4bit is still ~35GB before accounting for context space. Even a 4090 can't run a 70b. Supposedly Mistral NeMo better than Llama-3-8b, which is the more apt comparison, although benchmarks usually don't tell the full story; we'll see how it does on the LMSYS Chatbot Arena leaderboards. The other (huge) advantage of Mistral NeMo over Llama-3-8b is the massive context window: 128k (and supposedly 1MM with RoPE scaling, according to their HF repo), vs 8k. Also, this was trained with 8bit quantization awareness, so it should handle quantization better than the Llama 3 series in general, which will help more people be able to run it locally. You don't need a 4090.
- adt 2y agoThat's 3 releases for Mistral in 24 hours. https://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/
- ofermend 2y agoCongrats. Very exciting to see continued innovation around smaller models, that can perform much better than larger models. This enables faster inference and makes them more ubiquitous.
- andrethegiant 2y agoI still don’t understand the business model of releasing open source gen AI models. If this took 3072 H100s to train, why are they releasing it for free? I understand they charge people when renting from their platform, but why permit people to run it themselves?
- kaoD 2y ago> but why permit people to run it themselves? I wouldn't worry about that if I were them: it's been shown again and again that people will pay for convenience. What I'd worry about is Amazon/Cloudflare repackaging my model and outcompeting my platform.
- andrethegiant 2y ago> What I'd worry about is Amazon/Cloudflare repackaging my model and outcompeting my platform. Why let Amazon/Cloudflare repackage it?
- bilbo0s 2y agoHow would you stop them? The license is Apache 2.
- andrethegiant 2y agoThat's my question -- why license as Apache 2
- bilbo0s 2y agoWhat license would allow complete freedom for everyone else, but constrain Amazon and Cloudflare?
- supriyo-biswas 2y agoThe LLaMa license is a good start.
- davidzweig 2y agoDid anyone try to check how are it's multilingual skills vs. Gemma 2? On the page, it's compared with LLama 3 only.
- moffkalast 2y agoWell it's not on Le Chat, it's not on LMSys, it has a new tokenizer that breaks llama.cpp compatibility, and I'm sure as hell not gonna run it with Crapformers at 0.1x speed which as of right now seems to be the only way to actually test it out.
- I_am_tiberius 2y agoThe last time I tried a Mistral model, it didn't answer most of my questions, because of "policy" reasons. I hope they fixed that. OpenAI at least only tells me that it's a policy issue but still answers most of the time.
- zone411 2y agoInteresting that the benchmarks they show have it outperforming Gemma 2 9B and Llama 3 8B, but it does a lot worse on my NYT Connections benchmark (5.1 vs 16.3 and 12.3). The new GPT-4o mini also does better at 14.3. It's just one benchmark though, so looking forward to additional scores.
- chant4747 2y agoCan you help me understand why people seem to think of Connections as a more robust indicator of (general) performance than benchmarks typically used for eval? It seems to me that while the game is very challenging for people it’s not necessarily an indicator of generalization. I can see how it’s useful - but I have trouble seeing how a low score on it would indicate low performance on most tasks. Thanks and hopefully this isn’t perceived as offensive. Just trying to learn more about it. edit: I realize you yourself indicate that it's "just one benchmark" - I am more asking about the broader usage I have seen here on HN comments from several people.
- zone411 2y agoThe most interesting thing about it is that it’s the type of task where you'd expect LLMs to do well, yet the best models only score around 30%, while top humans get 100%. Many other benchmarks are also getting close to saturation.
- lostmsu 2y agoGonna wait for LMSYS benchmarks. The "standard" benchmarks all seem unreliable.
- wkcheng 2y agoDoes anyone know whether the 128K is input tokens only? There are a lot of models that have a large context window for input but a small output context. If this actually has 128k tokens shared between input and output, that would be a game changer.
- eigenvalue 2y agoI have to say, the experience of trying to sign up for Nvidia Enterprise so you can try the "NIM" packaged version of this model, is just icky and and awful now that I've gotten used to actually free and open models and software. It feels much nicer and more free to be able to clone llama.cpp and wget a .gguf model file from huggingface without any registration at all. Especially since it has now been several hours since I signed up for the Nvidia account and it still says on the website "Your License Should be Active Momentarily | We're setting up your credentials to download NIMs." I really don't get Nvidia's thinking with this. They basically have a hardware monopoly. I shelled out the $4,000 or so to buy two of their 4090 GPUs. Why are they still insisting on torturing me with jumping through these awful hoops? They should just be glad that they're winning and embrace freedom.
- pennomi 2y agoThis is what you get when managers design a software tool instead of engineers designing it.
- lopuhin 2y agoAlso I don't think you can use NIM packages in production without a subscription, and I wasn't able to find the cost without signing up. Also NIM package for Mistral Nemo is not yet available anyways.
- hislaziness 2y agoI just checked huggingface and the model files download is about 25GB but in a comment below someone mentioned it is 8fp quantized model. Trying to understand how the quantization affects the model (and RAM) size. Can someone please enlighten.
- frontierkodiak 2y agoSure. The talk about 8bit refers to quantization-aware training. Pretty common in image models these days to reduce the impact of quantization on accuracy. Typically this might mean that you simulate an 8bit forward pass to ensure that the model is robust to quantization ‘noise’. You still use FP16/32 for backward pass & weight updates for numerical stability. It’s just a way to optimize the model in anticipation of future quantization. The experience of using an 8-bit Nemo quant should more closely mirror that of using the full-fat bf16 model compared to if they hadn’t used QAT.
- PoignardAzur 2y ago> Mistral NeMo uses a new tokenizer, Tekken, based on Tiktoken, that was trained on over more than 100 languages, and compresses natural language text and source code more efficiently than the SentencePiece tokenizer used in previous Mistral models. From Mistral's page about Tekken: > Our newest tokenizer, tekken, uses the Byte-Pair Encoding (BPE) with Tiktoken. Does that mean that Mistral found that BPE is more efficient than unigram models? Because otherwise, I don't understand why AI companies keep using BPE for their token sets. Unigram methods leads to more legible tokens, fewer glitch tokens, fewer super-long outlier tokens, etc.
- danielhanchen 2y agoI just managed to make Mistral NeMo 4bit QLoRA finetuning fit in under 12GB, so it fits in a free Google Colab with a Tesla T4 GPU! VRAM is shaved by 60% and finetuning is also 2x faster! Colab: https://colab.research.google.com/github/unslothai/studio/blob/main/colabs/mistral_nemo_12b.ipynb https://colab.research.google.com/github/unslothai/studio/bl...