6 ms·
Mistral Medium 3.5
- amunozo 5mo agoI want to believe it's gonna be good, but after trying GPT-5.5 even the most advanced Chinese models seem depressing.
- ako 5mo agoThen you’ll be happy to learn it’s not Chinese
- dotancohen 5mo agoGP is stating that the second best in the field, the Chinese, is so far behind the best in the field, GPT 5.5, that it is not even worth testing anything else.
- amunozo 5mo agoThanks for the translation, I did not express it very clearly. Anything that I try is so much worse.
- Ritewut 5mo agoIs GPT 5.5 the best in the field? I think Opus is still better despite Anthropic's recent stumbling.
- amunozo 5mo agoI did not try much Opus recently as I had a Codex subscription and heard bad things, but Opus is super good too. Let's say compared to any of them.
- r0b05 5mo agoThis is a French model sir
- lava_pidgeon 5mo agoHonestly I depends on the context which this performance matters. Mistral is quiet cheap
- manishsharan 5mo agoI am not following this obsession with SOTA and benchmark rankings I have been using DeepSeek and GLMnmodels with OpenCode and Codex and Claudr side by side. I have not found the Chinese models lacking. I enjoy for coding and like to maintain full control of my codebade and deeply care about the GOF patterns. So I am very stringent in terms of what I want the LLM to code and how to code. So from my perspective, they are all about the same.
- amunozo 5mo agoThat I agree with, but for more complex autonomous changes the differences are considerable. However, it seems that most models will reach the saturation time in which they will be useful for almost everything and the difference will be in more and more niche and specialized tasks.
- InputName 5mo agoLooks at first graph. It's SWE-Bench Verified. A benchmark Open-AI stopped using two months ago due to contamination. Doesn't look to promising. Is there any reason to consider Mistral other than it's not US?
- tpurves 5mo agoIf it's not US and it's within a few percent of SOTA that might be good enough for a lot of people (eg Europeans)
- NitpickLawyer 5mo agoGemma has been better for us at EU languages than mistral (for comparable sized models) :/ so ... dunno. What mistral does well and others are lagging behind is deploying on prem with their engineers and know-how, offering tuned models for your tasks and finetuning on your own data. (I expect google to start offering this next)
- deaux 5mo agoIt's sad that despite their strength in this for onprem, they're so behind on this in the cloud. No publicly available cloud SFT at all. Meanwhile Google has been offering that for years - though remains to be seen if they will for Gemini 3 when GA. And on top of it a range of providers like Fireworks and so on that offer it for Chinese models. This seems such an obvious thing for Mistral to offer.
- amunozo 5mo agoPrice and speed.
- 2ndorderthought 5mo agoThey did not stop using it due to contamination. They said it's flawed and indirectly said anthropics results were impossible. It's very possible they are sore losers
- spwa4 5mo agoTLDR: Mistral Medium 3.5, text-only, 128B dense model, 256k context window, modified MIT license. Model is ~140G ... https://huggingface.co/mistralai/Mistral-Medium-3.5-128B https://huggingface.co/mistralai/Mistral-Medium-3.5-128B They more or less claim this exceeds Claude Sonnet 3.5 on most things, but is worse than Sonnet 3.6, and exceeds all other open models. Oh and they have a cloud service that will code your apps "in the cloud". But, yeah, at this point, so does my cat. And, yes, unsloth is on it: https://huggingface.co/unsloth/Mistral-Medium-3.5-128B-GGUF https://huggingface.co/unsloth/Mistral-Medium-3.5-128B-GGUF (but 4bit quant is 75G)
- Marciplan 5mo agoYou mean Sonnet 4.5 and 4.6 riight
- spwa4 5mo agoright
- wolttam 5mo agoSonnet 4.5 and 4.6* There is no way it exceeds “all other” open models - but it does exceed all of Mistral’s past models. You can see it getting blown past by GLM 5.1 and Kimi in this. Still excited to give it a try
- 2ndorderthought 5mo agoIt looks like qwen 3.6 is winning and smaller for the April small model roll out
- pama 5mo agoUnfortunately they only compare to old “all other open models”. There are probably over 10 other open models better than it by now.
- mtct88 5mo agoIt's okay, nothing exceptional, but any news from non US and non Chinese models is still good news.
- pb7 5mo agoThis is the bar for Europe, huh?
- saulapremium 5mo ago[flagged]
- pb7 5mo agoThe fact that this comment is still up hours later but my comment below participating in the discussion got flagged should tell one everything they need to know about the intellectual rigor here.
- saulapremium 5mo agoOh it's flagged as well, and I admit that it was low effort. But your comment served no purpose besides from provocation, and I guess it worked.
- amunozo 5mo agoThis is the bar for anybody that's not the frontier labs.
- locknitpicker 5mo ago> This is the bar for Europe, huh? A few months ago China was being criticized left and right on how somehow it was not able to compete, and once DeepSeek showed up then all the hatred shifted onto how China was actually competing but exploring unfair competitive advantages. Funny how that works. Also, aren't the likes of OpenAI burning through over $2 of investment for each $1 of revenue?
- wyre 5mo agoI'm rooting for Mistral. It seems they are making a big bet that smaller models will win over larger ones and I can see it happening. I was running some simple chat and tool-calling benchmarks for small models and Mistral Small 4 performed well for it's price ($.15/$.60). Seeing this today got me excited, benchmarks seems solid compared to models much larger, but it's priced higher than Haiku, 5.4 mini, and all the the Chinese models it's comparing itself too. It's not even winning those benches either, just being competitive with them, which is great, those models are 5x+ the size, but they are also 1/2 the price. Hard to be excited about that.
- postalcoder 5mo agoThis release Mistral really reminds you of the gap between the frontier labs and everyone else. Pre-agent, there wasn't always an obvious difference between models. Various models had their charms. Nowadays, I don't want to entertain anything less than the frontier models. The difference in capability is enormous and choosing anything less has a real cost in terms of productivity. I've been a big fan of the smaller labs like Mistral and especially Cohere but it's been a while since I've been excited by a release by either company. That said, I'm using mistral voxtral realtime daily – it's great.
- onlyrealcuzzo 5mo ago> Pre-agent, there wasn't always an obvious difference between models. Various models had their charms. Nowadays, I don't want to entertain anything less than the frontier models. The difference in capability is enormous and choosing anything less has a real cost in terms of productivity. It's just apples to oranges. There is not a clear, across the board, winner on non-agentic tasks between Gemini, ChatGPT, and Claude - the simple chatbot interface. But Claude Code is substantially better than Codex which itself is notably better than Gemini-cli. In this vein, it should not be surprising that Claude Code is way better than non-frontier models for agentic coding... It's substantially better than other frontier models at specialized agentic tasks.
- nothinkjustai 5mo agoCC is not better than Codex, nor is it better than OpenCode, Crush, Pi etc…
- postalcoder 5mo agoI think there's a fair amount of evidence that the heavy harnesses actually drag down performance compared to bare harnesses.
- philipbjorge 5mo agoI’ve been comparing Claude Code and Codex extensively side by side over the past couple of weeks with my favorite prompting framework superpowers… From my perspective, Claude Code is decidedly not better than Codex. They’re slightly different and work better together. I would have no issues dropping CC entirely and using codex 100%. If you’re working off of “defaults”, in other words no custom prompting, Claude Code does perform a lot better out of the box. I think this matters, but if you’re a professional software developer, I’d make the case that you should be owning your tools and moving beyond the baked in prompts.
- vessenes 5mo agoAs always, rooting for these guys — model and national diversity is great. This looks like a solid foundation to build on; hopefully the 3.6/3.7 will dial in more gains. It looks like maybe from the computer use benchmarks that their vision pipeline could use improvement, but that’s just speculation. The different results on some benchmarks vibes as if this is truly an independently trained model, not just exfiltrated frontier logs, which I think is also really important - having different weight architectures inside a particular model seems like a benefit on its own when viewed from a global systems architecture perspective.
- sayYayToLife 5mo ago[dead]
- Tepix 5mo agoI use Mistral Le Chat quite a bit. One thing in particular I was disappointed in was its bad explanations when asking about French grammar. It made multiple mistakes and the other models got it right, even Qwen 3.6 27b! Anyway, I'm hoping they catch up some more.
- kubb 5mo agoThere's a good chance that they'll catch up. The "AI race" is a race to the bottom, with the leaders blowing huge wads of cash on capabilities that get replicated months later by the competition at a fraction of the cost. The only benefit of leading is mindshare. OpenAI is doubling down on that, by investing in communication companies. That's their pathetic attempt at a "moat".
- pb7 5mo agoThey catch up by distilling frontier models. They will eventually figure out how to prevent that from happening. No one has any interest in investing tens of billions if the product can be copied and sold for less.
- amarcheschi 5mo ago>No one has any interest in investing tens of billions if the product can be copied and sold for less. That is what has happened until now though
- mark_l_watson 5mo agoI like the idea of Mistral, but the last time I evaluated Mistral Vibe it was really nice for $15/month but not as effective as Gemini Plus with AntiGravity and gemini-cli. I am currently running Gemini Ultra on a 3 month 'special deal' and AntiGravity with Opus 4.7 tokens is pretty much fantastic. That said, when I stop spending money on Gemini Ultra, I will give Mistral Vibe another 1-month test. I like the entire business model and vibe of Mistral so much more than OpenAI/Anthropic/Google but I also have stuff to get done. I am curious if Mistral Vibe for $15/month is a stable business model (i.e., can they make a profit).
- amunozo 5mo agoI'm testing it right now and it seems very buggy and unstable, just like before.
- danelski 5mo agoHow do you feel about the responsiveness of gemini-cli? I tried it on a paid plan and the 10-minute hang-ups (per step, not the whole plan execution) really break the illusion of performance gains, unless you run it in the background and do something else in the meantime. It's more noticeable when Americans are awake.
- mark_l_watson 5mo agoit is usually fast, but if gemini-cli or any other coding agent is sluggish I quit using it for a while.
- simonw 5mo agoI can't figure out if this is available in the official Mistral API or not. Their model listing API returns this: { "id": "mistral-medium-2508", "object": "model", "created": 1777479384, "owned_by": "mistralai", "capabilities": { "completion_chat": true, "function_calling": true, "reasoning": false, "completion_fim": false, "fine_tuning": true, "vision": true, "ocr": false, "classification": false, "moderation": false, "audio": false, "audio_transcription": false, "audio_transcription_realtime": false, "audio_speech": false }, "name": "mistral-medium-2508", "description": "Update on Mistral Medium 3 with improved capabilities.", "max_context_length": 131072, "aliases": [ "mistral-medium-latest", "mistral-medium", "mistral-vibe-cli-with-tools" ], "deprecation": null, "deprecation_replacement_model": null, "default_model_temperature": 0.3, "type": "base" }, So that has the alias "mistral-medium-latest", but the official ID is "mistral-medium-2508" which suggests it's the model they released in August 2025. But... that 1777479384 timestamp decodes to Wednesday, April 29, 2026 at 04:16:24 PM UTC So is that the new Mistral Medium?
- simonw 5mo agoSome poking around in the source code for https://github.com/mistralai/mistral-vibe https://github.com/mistralai/mistral-vibe got me to this: curl https://api.mistral.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $(llm keys get mistral)" \ -d '{ "model": "mistral-medium-3.5", "messages": [ {"role": "user", "content": "Generate an SVG of a pelican riding a bicycle"} ] }' Which did work: https://gist.github.com/simonw/f3158919b18d2c47863b0a5dc257a355#response https://gist.github.com/simonw/f3158919b18d2c47863b0a5dc257a... - it's pretty disappointing. Weird that it doesn't show up in the model list: curl https://api.mistral.ai/v1/models \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $(llm keys get mistral)" | jq
- Mashimo 5mo agoI also did some SVG tests, it's really bad. https://chat.mistral.ai/chat/897fbe7d-b1ae-4109-9b29-f3ccc4f8896d https://chat.mistral.ai/chat/897fbe7d-b1ae-4109-9b29-f3ccc4f...
- minimaxir 5mo agoIt's funny that 128B is now considered Medium. I remember back in the day when 355M parameters was considered medium with GPT-2.
- speedgoose 5mo agoAnd GPT-2 1.5B was considered too dangerous to release. They were perhaps right.
- Matl 5mo agoconsidered that by OpenAI for marketing purposes that is But yes, perhaps it would have been better for all of us if they haven't.
- refulgentis 5mo agoIn lockstep over the past month, a subset of people, un-labelable, unprompted, share this train of thought: - Mythos wasn't released widely. - But Anthropic shared info on it and said it was dangerous. - Anthropic is a company. - Companies like money. - Therefore Mythos is marketing hype. - Remember GPT-2? That also wasn't released. They said it was dangerous. - But, GPT-3, GPT-4, GPT-5, etc. were released. - Therefore GPT-2 being dangerous was marketing hype. I've seen the idea that GPT-2 not being released was marketing hype at least 6 times since Mythos was shared. It's Not Even Wrong, in the Pauli sense: they weren't selling anything! They weren't raising funding! What were they marketing!? And there's a lot more elided from history, ex. they didn't have an API yet. GPT-3 was released, a year or two later, and did have an API. But, no one used it, it wasn't good enough yet. And they did treat it as dangerous, it was wildly over-the-top manually monitored for anything resembling not-intended-use. I got permanently suspended for using the word "twink"
- Matl 5mo ago> I've seen the idea that GPT-2 not being released was marketing hype at least 6 times since Mythos was shared. That's not what I am saying. It's not that GPT-2 not being released was marketing hype, it's that OpenAI themselves claiming it's too dangerous to release specifically, implying it's close to AGI, (or something like that), was marketing hype.
- Giorgi 5mo agoOh they are still a thing?! Completely forgot about Mistral. I am assuming they are still burning trough investor money.
- sev_verso 5mo agoWhat's better than Voxtral for locally processed voice input? More competition is always better.
- danelski 5mo agoI believe they'll get profitable sooner than their frontier competition. Their operating costs seem to peanuts compared to the providers they're compared to most often while having the local advantage of not being Chinese nor American.
- kergonath 5mo ago> they are still burning trough investor money Difficult to say, this information is not really public. That said, those investors include EU agencies and European multinational companies and governments. It’s not as flashy as the ridiculous sums OpenAI is getting but it should be enough to keep them going for a while. They also have a different business model. They are selling their expertise to fine tune and adapt their models to on-premises computers (which they can help you build) to handle confidential data and information. I would not be surprised that the revenue they get from normal people is negligible in comparison.
- Giorgi 5mo agoOoh, ok so people got all worked up because it is EU vs USA thing.
- Havoc 5mo agoThink they’re positioned pretty well. They’ve got an edge in the European corporate space and don’t have ungodly large numbers to hit
- simjnd 5mo agoI'm not sure what people are on in the comments. It doesn't beat the other models, but it sure competes despite its size. GLM 5.1 is an excellent model, but even at Q4 you're looking at ~400GB. Kimi K2.5 is really good too, and at Q4 quantization you're looking at almost ~600GB. This model? You can run it at Q4 with 70GB of VRAM. This is approaching consumer level territory (you can get a Mac Studio with 128GB of RAM for ~3500 USD). For the Claude-pilled people, I don't know if you only run Opus but when I was on the Pro plan Sonnet was already extremely capable. This beats the latest Sonnet while running locally, without anyone charging you extra for having HERMES.md in your repo, or locking you out of your account on a whim. Mistral has never been competitive at the frontier, but maybe that is not what we need from them. Having Pareto models that get you 80% of the frontier at 20% of the cost/size sounds really good to me.
- redrove 5mo agoIt’s 128b dense model. Good luck getting more than 3t/s out of a mac. It doesn’t matter if it fits or not.
- zozbot234 5mo agoYou could run it on a single Mac Studio with M3 Ultra, or two Mac Studios with M4 Max at higher perf than that. And lightly quantizing this could give us modern dense models in the ~80GB size range, which is a very compelling target.
- freakynit 5mo agoWouldn't matter much still. M3 ultra has 819GB/s unified memory bandwidth. That means theoretical max tokem rate is 819/128 =~ 6.39 t/s. At 80 GB (5 bit quantization), its still near about 10 t/s ... far from a good coding experience. Also, these are theoretical max.. real world token generation rates would be at least 15-20% less.
- gregsadetsky 5mo agoI didn't know about HERMES.md ... (??) - found information here for others who are curious https://github.com/anthropics/claude-code/issues/53262 https://github.com/anthropics/claude-code/issues/53262
- Mashimo 5mo agoCompared to all other hosted LLMs that I have tested, Mistral seems to be the only one with rather strict CSP headers. When you ask them to create a website with some javascript library it will not preview, even though le chat offers canvas mode. Sometimes when a new release comes around from any provider I just want to test it a bit on the web. without paying and using an agent harness. Why are they like this ;_; Edit: Christ on a bike it's bad at drawing SVGs https://chat.mistral.ai/chat/23214adb-5530-4af9-bb47-90f52192f274 https://chat.mistral.ai/chat/23214adb-5530-4af9-bb47-90f5219...
- 2ndorderthought 5mo agoI have never wanted, needed or hoped to draw svgs with an LLM. All of the models suck at it, some are just more fun or something.
- Mashimo 5mo agoI can't speak for what you consider sucking, but there is a significant difference between Mistral and Kimi or Gemini. I find the others to be usable for my needs.
- 2ndorderthought 5mo agoI agree there is a difference but does that translate to anything? It's not the same operations used to write code, and it's kind of useless. I wouldn't waste my power bill ensuring a model I was releasing was good at it.
- Mashimo 5mo ago> It's not the same operations used to write code Is it not? It's html and javascript. And not even attempting to draw details that other models do. When I try other html / js prompts it also lacks behind china models from over half a year ago. I mean worse then GLM 4.7.
- 5mo ago
- deleted 5mo ago[deleted]
- seb_lz 5mo agoI'm using mistral-medium-2508 for some text transformation operations. It's giving me better results than mistral-large for my use cases. Looking forward to testing this new model, although I'm not sure if it's really meant at replacing the previous medium model since it's a lot more expensive and presented more as a coding / agentic model (mistral-medium-2508 was priced $0.4/$2 per 1M tokens, mistral-medium-3.5 is $1,5/$7.5).
- hulk-konen 5mo agoI actually use Mistral Large to go through some large text chunks (in production). It gives about the same level of results as Sonnet, while being 90% cheaper. Definitely wouldn't use it for coding, but for this text-analyzing task it has been great. Much better than all the latest Chinese models, for example. So I was waiting for this release and it's... 5x more expensive than the latest Mistral Large. So now I'm worried they'll pull the plug on the cheap Large when their releases roll over to that one.
- zozbot234 5mo agoWhy does this matter if the model is open? It can be offered by competitive third-party providers, there's no rug pull.
- hulk-konen 5mo agoRight now it's really not offered by third parties. I found it via a single provider (BitDeer). I'm not sure I'd trust them with my customers' data. Also, considering the model is getting a bit old, I wouldn't expect them to keep offering it forever. Anyhow, competition is fierce. I'll have some model I can use in the future, even if it's not dirt cheap like current Mistral Large is.
- barrell 5mo agoYeah I use Mistral Large for a lot of formatting work. For this one use case of mine, it outperforms frontier models by a significant margin. I've found tons of use cases for mistral small as well. I'd love to use Mistral for more tasks, but Mistral Large doesn't quite cut it for all tasks. So on the one hand, I'm excited there is another model, and presumably more performant based on the price? But the fact it's a "Medium" and 5x the price of the Large definitely concerns me. The entire release is also about Vibe Coding, and so I'm not even sure if this model is applicable outside of coding, or even worth testing.
- syntaxing 5mo agoThis is a very interesting strategy that might pay off. This model is a very good option for enterprise self host. I would argue a lot of companies are VRAM constrained rather than compute constrained. You could fit 4-5 running instances on one H100 cluster where you can only fit 1-2 Kimi K2 or GLM5.
- 2001zhaozhao 5mo agoThis is 128B dense though. the K/V cache on long context is going to be massive
- Havoc 5mo agoDon’t think kv size correlates to dense/moe
- zozbot234 5mo agoKV size correlates with attention parameters which are a subset of active parameters. So a typical MoE model will have way lower KV size than a dense model of equal total parameter count.
- syntaxing 5mo agoWith turbo quant, you would reduce it by over 6X.
- sayYayToLife 5mo ago[dead]
- maelito 5mo agoGiven what Vibe already did in the previous versions with codestral-v2, that's great news. Keep up the good work ! I don't want to depend on the world's two hungry superpowers.
- Alifatisk 5mo agoA 1000B model, can we call it 1KB model?
- schipperai 5mo agoWith most OSS releases being MoEs, and modern GPUs optimized for MoEs, can somebody with knowledge of the topic explain or speculate why Mistral might have opted for a dense model?
- ac29 5mo agoModern GPUs aren't optimized for MoEs though? The advantage to a dense model like this Mistral one is that it is as smart as a much larger MoE model so it can fit on less GPUs. The tradeoff is that it is much slower since it has to read 100% of its weights for every token, MoE models typically only read about a tenth (though sparsity levels vary).
- schipperai 5mo agoThanks, makes sense. I meant Blackwell is explicitly optimized for MoEs.
- andhuman 5mo agoThe Vibe CLI is really bad on Windows, sure they don’t officially support it, so can’t blame them, but a FYI for anyone wanting to try it. It can’t get find and replace right.
- deferredgrant 5mo ago[flagged]
- KronisLV 5mo agoFor it's size, that's really good! Though I bet it being a dense model probably helps a lot, if it was MoE at that size, I bet the benchmark performance would go quite a bit down (which consequently would also mean that I'd at least be able to run it with decent tokens/second, with the bunch of Nvidia L4 cards available to me, which presently are only okay with MoE models). It's cool that they added comparisons to their own Mistral Small 4 119B A7B, which kind of shows that! They could have also included comparisons to something like Qwen Coder Next 80B A3B (or maybe the newer Qwen 3.6 35B A3B, or the 27B dense one), maybe DeepSeek V4 Flash 284B A13B, or the older GPT-OSS 120B A5B to illustrate that difference and where their model sits even better, it would probably give a more positive picture than just comparing themselves against a bunch of bigger models! Come to think of it, alongside throwing some money at DeepSeek not just Anthropic, I probably should get a Mistral subscription as well sometime, to see how they perform on various tasks - cause they seem pretty cost effective and it's nice to support at least some EU orgs: https://mistral.ai/pricing https://mistral.ai/pricing
- jollymonATX 5mo agoOuch. Maybe they have a captive buying market to insulate them from actual market forces or ???
- antirez 5mo agoThe problem with this model is that DeepSeek v4 Flash runs quite well quantized to 2 bit (see https://github.com/antirez/llama.cpp-deepseek-v4-flash https://github.com/antirez/llama.cpp-deepseek-v4-flash), at 30 t/s generation and 400 t/s prefill in a M3 Ultra (and not too much slower on a 128GB MacBook Pro M3 Max). It works as a good coding agent with opencode/pi, tool calling is very reliable, and so forth. All this at a speed that a 120B dense model can never achieve. So it has to compete not just with models that fit 4-bit quantized the same size, but with an 86GB GGUF file of DeepSeek v4 Flash, and it is not very easy to win in practical terms for local inference. Note: I have more uncommitted speed improvements in my tree that I'll push soon, the current tree could be a little bit slower but not much, still super usable. I don't understand one thing about Mistral, which I'm a fan being in Europe: they opened the open weights MoE show with Mixtral. Why are they now releasing dense models of significant sizes? In this way you don't compete in any credible space, nor local inference, nor remote inference since the model is far from SOTA and not cheap to serve. So why they are training such dense big models? Dense models have a place in the few tens of billion parameters, as Qwen 3.6 27B shows, but if you go 5 times that, it is no longer a fit, unless you are crushing with capabilities anything requiring the same VRAM, which is not the case.
- zozbot234 5mo agoYour GitHub link only says "The model quantized in this way behaves very very well in the chat, frontier-model vibes, but it was not extensively tested." This is hardly relevant to how it behaves in agentic workflows, we're aware of how often they degrade severely with Q2 quantization. If this quantized Flash can keep up reasonable quality and performance at larger context lengths (which seems to be a key feature of the V4 series) it could be a very reasonable competitor to models in the same weight class like Qwen 3 Coder-Next 80B.
- antirez 5mo agoNope it works great with opencode as a agent, you can build a game or things like that. It works. The trick is a mix among the quantization I used, which is very asymmetric, and the fact that I guess DeepSeek v4 Flash tolerates extreme quantizations better than anything I saw in the past. What I used was up/gate of routed models, IQ2_XXS, out -> Q2_K, then I quantized routing, projections, shared experts to Q8. The trick is that the very sensible parts are a small amount of the weights, and they are kept very high quality.
- vicchenai 5mo ago[dead]
- vjay15 5mo agoMistral is playing the long game here ngl, lower sized models, lower costs, overall good enough performance!
- dznan 5mo agoI hack you
- Scroll_Swe 5mo agoLe Chat fast is really good now. Did a few normal queries and noticed and improvement. Very well done.