16 ms·
The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last ye
by cco 1y ago
The lede is being missed imo.
gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year.
I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two.
But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding error) on my laptop. No $200/month subscription, no lakes being drained, etc.
I'm blown away.
- MattSayar 1y agoWhat's your experience with the quality of LLMs running on your phone?
- NoDoo 1y agoI've run qwen3 4B on my phone, it's not the best but it's better than old gpt-3.5. It also does have a reasoning mode, and in reasoning mode it's better than the original gpt-4 and rhe original gpt-4o, but not the latest gpt-4o. I get usable speed, but it's not really comparable to most cloud hosted models.
- NoDoo 1y agoI'm on android so I've used termux+ollama, but if you don't want to set that up in a terminal or want a GUI pocketpal AI is a really good app for both android and iOS. It let's you run hugging face models.
- cco 1y agoAs other said, around gpt 3.5 level so three or four years behind SOTA today at reasonable (but not quick) speed.
- datadrivenangel 1y agoNow to embrace jevon's paradox and expand usage until we're back to draining lakes so that your agentic refrigerator can simulate sentience.
- herval 1y agoIn the future, your Samsung fridge will also need your AI girlfriend
- throw310822 1y agoIn the future, while you're away your Samsung fridge will use electricity to chat up the Whirlpool washing machine.
- pryelluw 1y agoIn Zap Brannigans voice: “I am well versed in the lost art form of delicates seduction.”
- hkt 1y agos/need/be/
- spauldo 1y ago"Now I've been admitted to Refrigerator Heaven..."
- nudgeOrnurture 1y agoand she will tell neither of you two who she got that cute little pixel badge from which will make you jealous and then the microwave will tell you it's been hustling on the side as a PAI and that it can get that info ... at the cost of a little upgrade
- turnsout 1y agoThe environmentalist in me loves the fact that LLM progress has mostly been focused on doing more with the same hardware, rather than horizontal scaling. I guess given GPU shortages that makes sense, but it really does feel like the value of my hardware (a laptop in my case) is going up over time, not down. Also, just wanted to credit you for being one of the five people on Earth who knows the correct spelling of "lede."
- twixfel 1y ago> Also, just wanted to credit you for being one of the five people on Earth who knows the correct spelling of "lede." Not in the UK it isn’t.
- turnsout 1y agoYes, it is, although it's primarily a US journalistic convention. "Lede" is a publishing industry word referring to the most important leading detail of a story. It's spelled intentionally "incorrectly" to disambiguate it from the metal lead, which was used in typesetting at the time.
- twixfel 1y agoNo, it isn't. It's an American thing at most and possibly also a false etymology. Its prime usage appears to be for people in HN threads to pat each other on the back for knowing it's "real" spelling.
- black3r 1y agocan you please give an estimate how much slower/faster is it on your macbook compared to comparable models running in the cloud?
- syntaxing 1y agoYou can get a pretty good estimate depending on your memory bandwidth. Too many parameters can change with local models (quantization, fast attention, etc). But the new models are MoE so they’re gonna be pretty fast.
- cco 1y agoSure. This is a thinking model, so I ran it against o4-mini, here are the results: * gpt-oss:20b * Time-to-first-token: 2.49 seconds * Time-to-completion: 51.47 seconds * Tokens-per-second: 2.19 * o4-mini on ChatGPT * Time-to-first-token: 2.50 seconds * Time-to-completion: 5.84 seconds * Tokens-per-second: 19.34 Time to first token was similar, but the thinking piece was _much_ faster on o4-mini. Thinking took the majority of the 51 seconds for gpt-oss:20b.
- parhamn 1y agoI just tested 120B from the Groq API on agentic stuff (multi-step function calling, similar to claude code) and it's not that good. Agentic fine-tuning seems key, hopefully someone drops one soon.
- AmazingTurtle 1y agoIm not sure if groq uses the proper harmony template?
- mathiaspoint 1y agoIt's really training not inference that drains the lakes.
- JKCalhoun 1y agoInteresting. I understand that, but I don't know to what degree. I mean the training, while expensive, is done once. The inference … besides being done by perhaps millions of clients, is done for, well, the life of the model anyway. Surely that adds up. It's hard to know, but I assume the user taking up the burden of the inference is perhaps doing so more efficiently? I mean, when I run a local model, it is plodding along — not as quick as the online model. So, slow and therefore I assume necessarily more power efficient.
- nudgeOrnurture 1y agoyou found a way to train only once until it "just works"?
- littlestymaar 1y agoTraining cost has increased a ton exactly because inference cost is the biggest problem: models are now trained on almost three orders of magnitude more data then what is compute-optimal to do (from the Chinchilla paper), because saving compute on inference makes it valuable to overtrain a smaller model to achieve similar performance for a bigger amount of training compute.
- syntaxing 1y agoInteresting, these models are better than the new Qwen releases?
- captainregex 1y agoI’m still trying to understand what is the biggest group of people that uses local AI (or will)? Students who don’t want to pay but somehow have the hardware? Devs who are price conscious and want free agentic coding? Local, in my experience, can’t even pull data from an image without hallucinating (Qwen 2.5 VI in that example). Hopefully local/small models keep getting better and devices get better at running bigger ones It feels like we do it because we can more than because it makes sense- which I am all for! I just wonder if i’m missing some kind of major use case all around me that justifies chaining together a bunch of mac studios or buying a really great graphics card. Tools like exo are cool and the idea of distributed compute is neat but what edge cases truly need it so badly that it’s worth all the effort?
- deleted 1y ago[deleted]
- canvascritic 1y agoHealthcare organizations that can't (easily) send data over the wire while remaining in compliance Organizations operating in high stakes environments Organizations with restrictive IT policies To name just a few -- well, the first two are special cases of the last one RE your hallucination concerns: the issue is overly broad ambitions. Local LLMs are not general purpose -- if what you want is local ChatGPT, you will have a bad time. You should have a highly focused use case, like "classify this free text as A or B" or "clean this up to conform to this standard": this is the sweet spot for a local model
- captainregex 1y agoAren’t there HIPPA compliant clouds? I thought Azure had an offer to that effect and I imagine that’s the type of place they’re doing a lot of things now. I’ve landed roughly where you have though- text stuff is fine but don’t ask it to interact with files/data you can’t copy paste into the box. If a user doesn’t care to go through the trouble to preserve privacy, and I think it’s fair to say a lot of people claim to care but their behavior doesn’t change, then I just don’t see it being a thing people bother with. Maybe something to use offline while on a plane? but even then I guess United will have Starlink soon so plane connectivity is gonna get better
- dongobread 1y agoHow up to date are you on current open weights models? After playing around with it for a few hours I find it to be nowhere near as good as Qwen3-30B-A3B. The world knowledge is severely lacking in particular.
- Nomadeon 1y agoAgree. Concrete example: "What was the Japanese codeword for Midway Island in WWII?" Answer on Wikipedia: https://en.wikipedia.org/wiki/Battle_of_Midway#U.S._code-breaking https://en.wikipedia.org/wiki/Battle_of_Midway#U.S._code-bre... dolphin3.0-llama3.1-8b Q4_K_S [4.69 GB on disk]: correct in <2 seconds deepseek-r1-0528-qwen3-8b Q6_K [6.73 GB]: correct in 10 seconds gpt-oss-20b MXFP4 [12.11 GB] low reasoning: wrong after 6 seconds gpt-oss-20b MXFP4 [12.11 GB] high reasoning: wrong after 3 minutes ! Yea yea it's only one question of nonsense trivia. I'm sure it was billions well spent. It's possible I'm using a poor temperature setting or something but since they weren't bothered enough to put it in the model card I'm not bothered to fuss with it.
- anorwell 1y agoI think your example reflects well on oss-20b, not poorly. It (may) show that they've been successful in separating reasoning from knowledge. You don't _want_ your small reasoning model to waste weights memorizing minutiae.
- deleted 1y ago[deleted]
- bigmanhank 1y agoNot true: During World War II the Imperial Japanese Navy referred to Midway Island in their communications as “Milano” (ミラノ). This was the official code word used when planning and executing operations against the island, including the Battle of Midway. 12.82 tok/sec 140 tokens 7.91s to first token openai/gpt-oss-20b
- WmWsjA6B29B4nfk 1y ago
- Cicero22 1y agoWhere did you get the top ten from? https://huggingface.co/spaces/TIGER-Lab/MMLU-Pro https://huggingface.co/spaces/TIGER-Lab/MMLU-Pro Are you discounting all of the self reported scores?
- zwischenzug 1y agoCame here to say this. It's behind the 14b Phi-reasoning-plus (which is self-reported). I don't understand why "TIGER-LAb"-sourced scores are 'unknown' in terms of model size?
- int_19h 1y agoI tried 20b locally and it couldn't reason a way out of a basic river crossing puzzle with labels changed. That is not anywhere near SOTA. In fact it's worse than many local models that can do it, including e.g. QwQ-32b.
- robwwilliams 1y agoWell river crossings are one type of problem. My real world problem is proofing and minor editing of text. A version installed on my portable would be great.
- cosmojg 1y agoHave you tried Google's Gemma-3n-E4B-IT in their AI Edge Gallery app? It's the first local model that's really blown me away with its power-to-speed ratio on a mobile device. See: https://github.com/google-ai-edge/gallery/releases/tag/1.0.3 https://github.com/google-ai-edge/gallery/releases/tag/1.0.3
- 1123581321 1y agoDozens of locally runnable models can already do that.
- golol 1y agoI heard the OSSmodels are terrible at anything other than math, code etc.
- mark_l_watson 1y agoYes, I always evaluate models on my own prompts and use cases. I glance at evaluation postings but I am also only interested in my own use cases.
- 9rx 1y agoI tried the two US presidents having the same parents one, and while it understood the intent, it got caught up in being adamant that Joe Biden won the election in 2024 and anything I do to try and tell it otherwise is dismissed as being false and expresses quite definitely that I need to do proper research with legitimate sources.
- bakies 1y agoon your phone?
- npn 1y agoIt is not a frontier model. It's only good for benchmarks. Tried some tasks and it is even worse than gemma 3n.
- snthpy 1y agoFor me the biggest benefit of open weights models is the ability to fine tune and adapt to different tasks.
- lend000 1y agoFor me the game changer here is the speed. On my local Mac I'm finally getting token counts that are faster than I can process the output (~96 tok/s), and the quality has been solid. I had previously tried some of the distilled qwen and deepseek models and they were just way too slow for me to seriously use them.
- decide1000 1y agoThe model is good and runs fine but if you want to be blown away again try Qwen3-30A-A3B-2507. It's 6gb bigger but the response is comparable or better and much faster to run. Gpt-oss-20B gives me 6 tok/sec while Qwen3 gives me 37 tok/sec. Qwen3 is not a reasoning model tho.
- raideno 1y agoHow much ram is in your Macbook Air M3 ? I have the 16Gb version and i was wondering whether i'll be able to run it or not.
- SergeAx 1y agoDid you mean "120b"? I am running 20b model locally right now, and it is pretty mediocre. Nothing near Gemini 2.5 Pro, which is my daily driver.
- benreesman 1y agoYou're going to freak out when you try the Chinese ones :)
- vonneumannstan 1y ago>no lakes being drained When you imagine a lake being drained to cool a datacenter do you ever consider where the water used for cooling goes? Do you imagine it disappears?
- nudgeOrnurture 1y agonot if the winds of fortune don't change;--but the weather, man, it's been getting less foreseeable than I was once used to
- jwr 1y agogpt-oss:20b is the best performing model on my spam filtering benchmarks (I wrote a despammer that uses an LLM). These are the simplified results (total percentage of correctly classified E-mails on both spam and ham testing data): gpt-oss:20b 95.6% gemma3:27b-it-qat 94.3% mistral-small3.2:24b-instruct-2506-q4_K_M 93.7% mistral-small3.2:24b-instruct-2506-q8_0 92.5% qwen3:32b-q4_K_M 89.2% qwen3:30b-a3b-q4_K_M 87.9% gemma3n:e4b-it-q4_K_M 84.9% deepseek-r1:8b 75.2% qwen3:30b-a3b-instruct-2507-q4_K_M 73.0% I'm quite happy, because it's also smaller and faster than gemma3.
- animanoir 1y ago[dead]
- latexr 1y agoI tried their live demo. It suggests three prompts, one of them being “How many R’s are in strawberry?” So I clicked that, and it answered there are three! I tried it thrice with the same result. It suggested the prompt. It’s infamous because models often get it wrong, they know it, and still they confidently suggested it and got it wrong.
- latexr 1y agoObviously I made a typo above. “Three” is the right answer, I meant that it answered there are “two” (the wrong answer). https://i.imgur.com/DgAvbee.png https://i.imgur.com/DgAvbee.png