7 ms·
I just got qwen 3.8 27b mlx running on my Macbook Pro and honestly I’m pretty blown away by how not-dumb it is.
by jonplackett 2mo ago
I just got qwen 3.8 27b mlx running on my Macbook Pro and honestly I’m pretty blown away by how not-dumb it is.
- rurban 2mo agoWe were trying running a local gpt-oss 80GB model on a H100, and honestly I was surprised how dumb it was.
- nick_ 2mo agogpt-oss is about a year older than qwen 3.8 27b
- anon373839 2mo agoGPT-OSS 20B didn’t really merit the fanfare even when it was released; it’s definitely not competitive now. Even the 120B version has been well eclipsed by smaller LLMs at this point. The last version of Qwen 27B/35B was better, and now the new one is even better than that!
- vikramkr 2mo agoWas there a more recent refresh or is this the model from a year ago? The frontier models were barely functional and almost useless a year ago (gpt oss was pre opus 4.5!) - I would be very surprised if the original drop is anything more than totally obsolete/irrelevant at this point
- rurban 2mo agoYes, the old entirely stupid old gpt-oss. But Sonnet and GPT were very useful then already, qwen also.
- vikramkr 2mo agoI find that I remember models being a lot better than they were, even when I remember them being not very good - because of a novelty factor ("whoa it can do that now?") mostly. And then I go back and look at them and its like, what how did I find this impressive. A funny example - I remember thinking "yeah sonnet 3.5 is a really good coding model" https://stack.convex.dev/using-cursor-claude-and-convex-to-build-a-social-media-scheduling-app https://stack.convex.dev/using-cursor-claude-and-convex-to-b... >Prompting Cursor to Scaffold my App: FAIL This was my first hurdle. >It became immediately apparent that I would not be able to prompt my way through the entire process. >While the tooling we have is undeniably powerful, it's not yet capable of completing most nontrivial tasks It couldn't run pnpm install lmao. Opus 4.5 was a crazy jump
- prettyblocks 2mo agoMy problem is how hot they run. I'm on an m4 pro. Do you have the same issue?
- downrightmike 2mo agoMineral oil bath?
- deleted 2mo ago[deleted]
- datadrivenangel 2mo agojust decent air cooling and you'll be okay. it will get up to 75/80C though for my M5 MBP.
- lukan 2mo agoI don't have the hardware but a often mentioned advice is to put your mac into energy saving mode - it still will work, a bit slower, but stays cool.
- PalmPilotProMax 2mo agoImagine spending all that money on Apple hardware only to throttle it to a fraction of its performance lol Steve Jobs would be proud. People really are holding their Apple hardware wrong.
- deleted 2mo ago[deleted]
- jonplackett 2mo agoIt’s hot and also LOUD and runs the battery down quick. But I’m having a lot of luck just running things when I’m away from the computer and can leave it plugged in. It starts going weird (unreliable and slow) with context over 80k so you have to pick tasks one at a time and baby sit a lot more than Claude. But it really is very capable and feels like there’s an intelligence there to talk to. Maybe gpt-4 level clever? I have an m5 max 64gb and I think anything slower would be quite painful.
- StarlaAtNight 2mo agohow quick does it respond? what are specs of your laptop?
- chorlton2080 2mo agoDoes it need to respond fast? For important applications, I'm sure we'd all be fine waiting 20 minutes for a high quality, usable answer. Or is it the need for interative refinements that make speed relevant?
- jonplackett 2mo agoIt requires patience but it’s more like waiting 5 mins for it to do tasks. You need to be much more involved though and do things slower than Claude where you can trust it to do a lot of tasks at once. It doesn’t have the context for that
- dominotw 2mo agoif you are so sure about what the final shape of your output is then its prbly not a common use of ai
- reverius42 2mo agoIf you are sure about the final shape of your output it's a great use case for AI, as you can define what you want in your prompt and refine towards it! It's where you don't know the end state you're looking for that you'll end up generating slop on top of slop and creating a whole Gastown just to power your Gastown.
- Gareth321 2mo agoI tried it on my M1 MacBook Pro. It's slow but surprisingly smart as a general purpose LLM. Maybe GPT-5.3 level. I gave it a bunch of tools and it can search the internet, make product recommendations, document, code, etc.
- alexpotato 2mo ago
- alexchantavy 2mo agoHow many tok/s are you getting? What gen mbp?
- dominotw 2mo agoi suspect ppl dropping generic "its awesome" comments are not actually using it and prbly just managed to get it running for a prompt or two.
- petcat 2mo agoYeah, that's my experience. It's a big "wow" factor to get a non-trivial LLM running on my Mac, but it's actually not that useful. Like trying to use Photoshop at 8 FPS.
- coldtea 2mo agoRegarding this analogy, fps don't matter as much for Photoshop, since it's not an immediate mode GUI. 8 fps would be quite ok for comfortably getting feedback on live image filters and such.
- LeBit 2mo agoIt’s not because you didn’t find use cases for local LLMs that there are none. I use local LLMs on my Mac Mini M4 Pro with 48G to review text messages tone, act as a text correction tool, act as a code review tool, to do code agent work, generate code snippets, etc Gemma 4 26B A4B gives me steady 20 tps.
- FireCrack 2mo agoI feel like it's 50/50 between people doing that, and people that have spent a lot of time tuning a system they are pointing at focused and well specified problems.
- mistersquid 2mo agoSeems threads about local LLMs on Apple hardware feature comments listing M3/4/5 at 48GB 64GB and not 128GB. That is, users with M-series hardware that have less-than-max RAM share results whereas users with max RAM do not. Speculating (not extrapolating), maybe users with machine that have max RAM are less interested in running local LLMs and are less averse to paying services for compute? Personally, I’d love to see what output max RAM M-series Apple hardware in these threads.
- applicative 2mo agoDid you read even the title?
- system2 2mo agoReread what he said maybe?
- riddlemethat 2mo agoI got the qwen 3.8 abliterated model running on my MacBook Pro M5 48GB and it's pretty nice having a local model that can do a lot of experimentation without rails.
- rahimnathwani 2mo agoorcarouter or obliteratus?
- tharkun__ 2mo agoIt was actually great. I have like a non-AI box so to speak 8GB VRAM, co-incidentally from a gaming PC ... All the previous models that were "frontier level, just try it!" but wouldn't run at all in agentic mode, including previous Qwens, just disappointed, period. Then I ran then Qwen 3.8 27b and while it was super slow (4t/s) it literally one-shotted creating a usable "web search/pull" skill for `pi.dev`. while any other model previously just entirely failed to create anything usable even with actual guidance. Since then I have actually gotten a gemma-4 12B qat 4bit quantized with a ~250MB MTP from unsloth to work with a 32k context "working" on this setup at 80-120 t/s. That's usable for private stuff on a co-incidental box! It's still only 32k context and it's entirely dumb vs. our API paid at-work Claude Opus. But for entirely private local stuff it's totally workable without breaking the bank even after all these AI price hikes!. I bought this rig literally just for gaming a month ago.
- aktenlage 2mo agoHave you tried a mixture of experts model? Dense models have been quite slow for me, as I have only 6 GB VRAM. But with llama.cpp and --cpu-moe I get 200 t/s input and almost 30 t/s output with Gemma 4 26B A3B, which feels ok to use. Would be interested about your mileage there.
- b112 2mo agoI wish qwen3.8 had a MoE variant, but the skinny is it won't be coming.
- tharkun__ 2mo agoIf I use the 12B Unified (dense) model I mentioned without MTP, then I get 37t/s, input ~700t/s. It's all still quite frustrating in the end, like a Claude from a very long time ago by now but usable. If I want 64k context, I can't use MTP. I still haven't decided whether I'd rather have 37t/s but it's "less dumb" or I want MTP speed but it's going off the rails more. All of this is also with `-ctv q4_0 -ctk q4_0)`, which is not ideal. I'm actually right now contending with 35k context but using q8_0 KV quantization. More like 35t/s coz with those settings I can't use MTP. But I'm not ready to go back to 4t/s. It's not interactive enough for me. That said, I had tried to use the Gemma E4B for example to have it build itself that websearch/fetch skill. It utterly failed, as did previous qwens. I don't see a Gemma 4 26B A3B GGUF for download, but there is a gemma-4-26B-A4B-it-MXFP4_MOE.gguf that should fit into my overall RAM and then use lots of CPU like the Qwen 3.8. I guess I'll give it a try just to see the difference in speed though I don't expect anything "usable" out of that tbh.
- velcrovan 2mo agoThat's funny, I downloaded the same model on my 48GB M4 Pro and gave it a problem to solve in an existing codebase, it spun its wheels for twenty minutes and then fell over dead. This was using LMStudio and pi as a harness; I never use pi for anything else, so maybe I'm holding it wrong.
- NamlchakKhandro 2mo ago[flagged]
- w-ll 2mo agoits all still somewhat of a dice roll
- cellularmitosis 2mo agoWe don’t know what quantization level was used for the weights or the kv cache for you or for parent poster, so this is probably an apples to oranges comparison.
- spacebacon 2mo ago[dead]
- s1gsegv 2mo agoThey made a kind of strange decision with Qwen3.8 27B, the template defaults the reasoning_effort to xhigh. I found if you set it to medium it doesn’t just sit there churning forever.
- ekianjo 2mo agoxhigh gives better results
- dofm 2mo agoNot necessarily. I have seen xhigh go down several rabbit holes, dwell on edge cases and write worse code as a result; it literally distracted itself into writing a complex chain of functions ignoring my prompt, when on “low” reasoning it gets it right on a prompt that requires a few lines of code in the right places. Simon Willison’s blog has another example (SVG of a circle). It’s a bit like how giving LLMs access to web search tools can cause them to go down a blind alley based on their first “reasoning” output that then leaves them unable to solve a puzzle correctly that they can fully solve on their own.
- hosteur 2mo agoHow much RAM? And what do you use it for if I might ask?
- bmitc 2mo agoWhich exact model are you running? With only 48GB of RAM, by the time I got a model small enough, it was pretty bad in performance both in speed and reasoning.