7 ms·
Bonsai 27B: A 27B-Class model that runs on a phone
- VaporJournalAPP 3mo ago[flagged]
- alvatech 3mo agoTIL that 1 bit models are actually 1.58 bit with three values +1, 0 and -1
- bensyverson 3mo agoYeah, it's an unfortunate convention from the very first "1 bit" model. But to be clear, Bonsai comes in both ternary and actual 1-bit variants.
- NitpickLawyer 3mo agoThere's two variants of this (or, as the joke goes, for very big values of bit): Ternary Bonsai 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, giving a true 1.71 effective bits per weight. 1-bit Bonsai 27B uses binary {−1, +1} weights with the same group-wise scaling, giving 1.125 effective bits per weight.
- PcChip 3mo agothis is a really dumb question, but how is -1 represented? is it a float? if so, how many bits is the float? I've never heard of a bit ever having more than two possible values
- zawaideh 3mo agoIt’s still a bit with only two possible values. But they add a scaling factor to a group of them (128 for example) which when you factor in, results in a fractional number of bits per parameter.
- throwawayffffas 3mo agoI believe the scaling comes in later, to turn the 1 and -1 into large numbers that may or may not activate the next layer. The way they do it is packing like the other comment says. Each byte represents 5 trinary values instead of 8 binary, and there is a little bit of waste.
- petu 3mo agopacking multiple trits together e.g. 5 trits (243 states) into a byte gives 1.6 bits per trit: https://compilade.net/blog/ternary-packing https://compilade.net/blog/ternary-packing
- JoshTriplett 3mo agoIt's impressive how close to optimal this is. You can beat the efficiency of 5 trits in 8 bits (1.6) with as few as 17 trits in 27 bits (~1.588), but once you account for rounding up to a whole number of bytes for practical reasons, then beating the efficiency requires going to at least 111 trits in 176 bits (~1.586), or perhaps more practically for fast unpacking, 161 trits in 256 bits (~1.59). At that level, even if you have, say, 27B trits, the more efficient encodings would save something like 38-45MB (theoretical limit ~48MB), likely at the cost of some slowdown.
- edflsafoiewq 3mo agoIt appears they are using Q2_0 in llama.cpp, which is 2 bits per weight + 1 float16 scale per group of 64 weights. This is inefficient in two ways: one bit pattern is wasted on each weight, since ternary weights only use {-1,0,1} and Q2_0 allows {-1,0,1,2}; and their group size is 128 weights, so the scale will be stored twice in two groups of 64 instead of stored only once in one group of 128. Their fork corrects the second inefficiency by using a group size of 128, but still uses 2-bit weights AFAICT. It's possible to pack 5 trits into a byte, but the unpacking is not very efficient. Another recent idea is to add the constraint that exactly one weight in each group of four be zero, which gives exactly 32 possible states, so it fits in 5 bits.
- snthpy 3mo agoThanks. This relates to some questions I've got. I was playing around with the previous generation smaller models and found that i wasn't getting any speedups from the T1 and T2 binary/ternary models compared to standard Q4 quants of straight qwen3.6 models. I was wondering whether unpacking of the ternary encoding was impacting inference speed? If that's the case then why not just train at Q2? I guess the counterargument is that then you lose the nice properties of things like the FairyFuse kernels. I wish there were some good discussions of these trade off.
- lioeters 3mo ago> never heard of a bit ever having more than two possible values It's not represented by a "bit", binary digit with value of 0 or 1; but with a "trit", ternary digit with value of {−1, 0, +1}.
- Dwedit 3mo ago1.6 bits if you want the most practical way to pack five 3-state numbers into a single byte. But even then, they usually pack four 4-state numbers instead.
- liuliu 3mo agoThe problem, of course, is if you run the UD_Q2 variant (Unsloth) which does only post-training, the number is pretty close to 1-bit model here and the 5% drop in tool-call is significant than it suggests in real-life use cases.
- liuliu 3mo agoYou also need to pay close attention to BFCLv3 multi-turn result, that helps you to get a sense how frequently these quants will be in a doom loop.
- h14h 3mo agoI'm curious what kind of results one could get from combining the clever quantization PrismML is doing here with something like LiquidAI's antidoom: https://github.com/Liquid4All/antidoom https://github.com/Liquid4All/antidoom
- ai_fry_ur_brain 3mo ago[dead]
- Havoc 3mo agoThis must be some sort of unpublished app? I can just see their image tool on the app store
- Catloafdev 3mo agoIt's a LLM model, not a phone app. Available on HuggingFace: https://huggingface.co/collections/prism-ml/bonsai-27b https://huggingface.co/collections/prism-ml/bonsai-27b
- Havoc 3mo agoIndeed. The article is about running it on a phone though, and shows an app with their branding running this in text mode on a phone. I'm asking where can I find this app to try what is being demonstrated in this article & video? Appstore only has an image gen app by them and other MLX apps I've tried don't seem to support this model
- smallerize 3mo agoOne of the links on the sidebar goes to "Locally AI" https://apps.apple.com/us/app/locally-ai-by-lm-studio/id6741426692 https://apps.apple.com/us/app/locally-ai-by-lm-studio/id6741... it requires an iPhone 17 Pro or Pro Max to run the 27B model though.
- Havoc 3mo agoI've got both that app and a 17 Pro, but it only lists one of the older Bonsai models not the 27B for me
- smallerize 3mo agoI only have an iPhone 14 Pro, but under "manage models" it's showing Bonsai 8B and Ternary Bonsai 8B.
- Havoc 3mo ago
- simonw 3mo agoThe models themselves are showing up on Hugging Face here: https://huggingface.co/prism-ml/models https://huggingface.co/prism-ml/models I've tried a couple in LM Studio - the GGUF one and the MLX one - but neither worked there. Anyone else get them to work? Might be that LM Studio needs to upgrade their llama.cpp or MLX engines first.
- trollbridge 3mo agoDidn't work for me in Unsloth, but it will probably be fixed in a day or two when the next batch of updates comes out.
- dofm 3mo agoHave Prism-ML upstreamed their forked code?
- bansaltushar 3mo agoDepending on which model you're running, you might need to use the custom forks. Details are here -> https://github.com/PrismML-Eng/Bonsai-demo/blob/main/README.md#upstream-status-for-ternary https://github.com/PrismML-Eng/Bonsai-demo/blob/main/README....
- motbus3 3mo agoI spent quite sometime trying to install their tools and nothing really worked. I used these repos you shared but the dependencies all fail on mac
- xyzsparetimexyz 3mo agoThat's awesome. What's the largest model that could fit onto a single 16gb gpu at 1.125 effects bits per weight?
- Catloafdev 3mo agoDoing some naive math, the F16 filesize is ~53.8gb, the 1-bit version is ~3.8gb, about 7% of the original size. The F16 size is roughly 2x param count, so that gives a rough ballpark of ~110B.
- kroaton 3mo agoWhich would be very interesting to test, as larger models (such as Deepseek V4 Flash or Qwen 397B) seem to compress better. Their Q2 quants are usable as is, even without the ternary compression.
- drob518 3mo agoYep, that’s the question. I asked just that when Bonsai’s first models got released. Super interesting if we can push the parameter count over 100B with 1.125 bit quantization and still keep pretty good performance versus 16-bit 100B models. That’s a definite sweet spot.
- syntaxing 3mo agoFor those curious about their demo, I’m pretty sure it’s using Locally AI (iOS only) that lmstudio acquired/aquihired a couple months ago.
- erelong 3mo agoI was trying Ornith 9B locally (it's up on Ollama) which claims: > Ornith-1.0-9B, which can be easily deployed on edge devices, matches or exceeds the performance of much larger models such as Gemma 4-31B and Qwen 3.6 35B. https://deep-reinforce.com/ornith_1_0.html https://deep-reinforce.com/ornith_1_0.html Only tried it so much so far; it did a little better than Qwen 9B
- liuliu 3mo agoNote that 3.5 9B cannot do thinking (while 3.6 27B can, pretty effectively, quite verbosely).
- gunalx 3mo ago3.5 9B can do thinking. Its just disabled by default in its gguf chat template.
- janalsncm 3mo agoIs that a 1-bit LLM? I don’t understand the connection with this article.
- erelong 3mo agoOh, I don't actually know the difference if you want to explain it The title says it's 27B grade running on a phone and what I was comparing it to in my mind was a model that runs at 35B grade that could presumably run on a phone "better"? edit: I asked AI for the difference and understand a little better, thanks for the heads up to learn the difference between models... I think the thing was, although ornith was created for a specific agentic purpose, it was still outperforming a previous generalist model I had running locally (so in my mind I thought it was still a better local model) - I'd like to try bonsai out if I can figure out how to run it lol
- syntaxing 3mo agoI don’t know if the llama cpp implementation is wonky (and only supports the binary version) but it’s a lot slower than 35B-A3B @ Q4_KM + MTP with CPU offloading.
- sigbottle 3mo agoWhat's the hiring space and business strategy around all of these smaller AI labs? Its really cool that people like these guys get paid to optimize models and give them out for free (open source). Do a lot of these labs have forward deployed engineers doing integrations with customers who want local models? Is there a general shift towards the local model crowd?
- trollbridge 3mo agoIf you read to the bottom of the page, it says they're funded by a few people, and one of them is Samsung. I'm betting Samsung wants to be able to ship a capable AI system on a future model of their phone so they can compete with Apple.
- doctoboggan 3mo agoAgreed, and the prevailing wisdom now seems to be that unless you can release a truly frontier model, you might as well release yours as open source to undercut your competition.
- ndrwdvvs 3mo agoor smart fridge
- try-working 3mo agoopen source is a GTM strategy
- thomasjb 3mo agoI've been watching and waiting for this, interested to see how smart it is, as it fits with my interest of getting the smartest possible model running in 10GB of VRAM (RTX3060 that has to drive 2 monitors and run an llm)
- dakolli 3mo agostart saving your money.
- thomasjb 3mo agoOr watch and wait as models get denser
- kennywinker 3mo agoToss the rtx into a cheapo optiplex or thinkcenter, and run it headless - the load on your machine is gonna make doing other stuff while it’s running painful. Plus that frees up the rest of your vram.
- xyzsparetimexyz 3mo agoWhy would they do that when they already have a perfectly good PC?
- kennywinker 3mo ago…? I said why. > the load on your machine is gonna make doing other stuff while it’s running painful Is your question about something else?
- thomasjb 3mo agoI aspire to someday move up to an AM5 system, but for now, the Dell T3610 has to do everything. Might get a second GPU for it though.
- HalfCrimp 3mo agoRunning an llm on a PC running a desktop environment really isn't that bad. You lose a bit of ram & vram but that only matters if you _reaaaally_ want to push to the max model size your hardware can handle. The biggest issue I've found is absent mindedly opening YouTube or the like that spike ram requirements and freezing the system up. But that's a me problem
- kristianp 3mo agoApparently Apple is "in talks" with the PrismML: https://www.cnbc.com/2026/07/14/apple-prismml-ai-compression-iphone.html https://www.cnbc.com/2026/07/14/apple-prismml-ai-compression...
- CharlesW 3mo agoNotably, PrismML CEO Babak Hassibi told CNBC this, so it’s either (1) bullshit, or (2) he just ended any chance of a relationship by leaking news of the talks.
- rogerkirkness 3mo agoApple would punish him severely unless they cleared it in advance, it might be to their advantage for some reason (negotiating with Google for Gemma rights? idk).
- segmondy 3mo agoApple is too desperate to be making demands. You confuse the Apple of yesterday and of today. Things change fast and things have changed.
- conception 3mo agoYou’re right, Apple only has 68 billion dollars in cash, up 40% since last year. Definitely on their last legs.
- brcmthrowaway 3mo agoI remember the infographics of 5-10 years ago. They had 200 billion. What happened? Sinking ship?
- malshe 3mo agoThey never had $200 billion in cash in any quarter. Not at least since 2014. Besides, how much cash to hold is an executive decision. If the cash balance went down, it doesn't mean their business is not doing well. Cash can be used for Capex, like what the hyperscalers are doing right now. It can be used for buying back equity.
- luckystarr 3mo agoTried it on Android and got "!!!!!!!!!!!!!" for answers.
- gunalx 3mo agoThe qwen models really seem to have this as a failure mode, its so annoying having a proper trace ending up in !!!!!! Garbage.
- amelius 3mo agoWait in a regular sentence, what is the probability of "!!!" being followed by "!"? Sounds like the model is not following a proper probabilistic choice here, so maybe more a programming error than a model training error.
- jldugger 3mo agoAfter the third !, the probability of a fourth probably skyrockets =)
- aesthesia 3mo agoOh, interesting. "!" is token id 0 in the Qwen tokenizer; I wonder if there's some tokenizer shenanigans either in inference or training that end up causing this specific behavior.
- verdverm 3mo agoThat's what happens when you quant too hard. I'm working on quant strats and evals for the same underlying qwen 27b models. When I saw 27b on a phone, I thought not fitting, big phone, or aggressive quant. NVFP4 still takes 27G before KV cache.
- erwan577 3mo agoThe KV-cache memory usage also seems remarkably frugal, even at the full context length. That could make this model particularly useful in multi-agent coding workflows. I wish KV-cache memory usage and related optimizations were discussed more clearly in new model announcements and demos.
- verdverm 3mo agoquanting kv cache hurts attention / recall, and long-form tasks by proxy. Model families and sizes have different tolerances to quant ting different parts of the model, same for intended tasks.
- pdfops 3mo ago[dead]
- wy35 3mo agoEntire blog post seems to be AI-generated :/
- wmf 3mo agoDo you think people who work on AI for a living are not going to use it?
- wy35 3mo agoOf course not, personally almost all of my code these days is generated. The LLM style of writing is just very distracting to read. “It unlocks X”, “Y changes the equation”, and why is there always something shifting? Makes my eyes glaze over in an otherwise interesting post.
- arjie 3mo agoThe text is mostly content-free. Headline + charts are enough for most HN stories.
- aesthesia 3mo agoIDK, Anthropic's public posts are mostly free of Claude-isms. (Their documentation less so...)
- theLiminator 3mo agoThis is useful research, but this particular model itself is likely absolutely useless.
- oceansweep 3mo agoWhy make this comment without having tried it first? It very clearly is not useless and performs a lot better than one might expect. I am currently waiting to do more benchmarks of it in comparison to the full weight model, but it seems promising/better than Mistral Nemo at a lower file size.
- Onavo 3mo agoI think what OP means is that the "minimum viable product" for a daily use LLM is probably somewhere around e.g. GPT 4o's level of intelligence (YMMV). Below a certain threshold, you are better off using specialized machine learning models rather than general purpose LLMs. It's very difficult to get that level of intelligence fully local on a mobile device without streaming to the cloud.
- mchusma 3mo agoI do think this would be interesting if they made these easy to finetune, as I do think this level of intelligence is likely sufficient for many applications and could be extremely cheap to run.
- drob518 3mo agoEvidence?
- Arcuru 3mo agoAwesome! I've been waiting for them to start scaling ternary models for over a year[1]. Excited to try it out, typical Qwen 27B is too heavy for me to run on my local hardware at reasonable speeds. [1] https://jackson.dev/post/dont-sleep-on-bitnet/ https://jackson.dev/post/dont-sleep-on-bitnet/
- drob518 3mo agoSame here. I’m excited to have a model that might be usable on a 16 GB laptop.
- comandillos 3mo agoQuite weird that heavy quantization method on a dense model gives better results than slightly quantized MoE models like 35B-A3B from Google. At this point all the different quantization and 'compression' (look at MPO applied to LLMs...) techniques start feeling a bit like snake oil. It's just gut feeling - or scores on benchmarks models are optimized for - what ends up deciding whether a technique is good enough or not.
- LeBit 3mo ago26B-A4B?
- motbus3 3mo agoI need help understanding this. I understood that the magic here is the quantization that allows it to use from 50G to 4G and their process retain most of the intelligence within Pareto limits of gain. And then they proceed to compare with other quantized models as in the level of intelligence per size. It gets to my attention though that the performance in tool calling is mostly affected which is a problem for other small models. How does this model compare to a recent 4G model? How do we know it retained intelligence from the parent rather then being fine tuned for the benchmarks? I am not shtng on them or anything. I'd rather find it amazing, BUT given my limited knowledge, I feel the results miss fair comparison plots and the ones might be misleading. Buy I also reckon it might be me the problem. Anyone care to explain this poor silly fellow some of those points?
- deleted 3mo ago[deleted]
- 0c3ca83 3mo ago[flagged]
- jedbrooke 3mo agofrom what I understand prismml isn’t doing a quant like normal models where you take a model trained at fp16 and then chop off some bits to reduce vram, but rather they’re training the model natively with 1 bit weights. It’s explained more in the article. They’re also doing some other tricks like a fp16 weight per block of 128 1bit weights to get some more data out of 1 bit weights
- parsimo2010 3mo agoThey aren’t training at all. They are quantizing existing models, it’s just that the process is different. The 27B uses Qwen3.6 27B as the base model.
- mehmetkose 3mo ago1-bit weights are not pointers so cpus can process them, storing them takes less space etc. tons of gains
- 0xbadcafebee 3mo ago27B is way more than you need for a phone. Doesn't matter how much you try to compress it, it's the wrong application of the wrong tool. There are already useful tiny models that fit on phones and do basic things really well. Dumb down a big model too much and it becomes worse than a small fine-tuned model.
- kamranjon 3mo agoAfter using a highly capable 2-bit quant as my daily driver for months now, I get pretty excited about releases like this. After a few days for the kinks to be worked out, I’ll be excited to try it.
- digdugdirk 3mo agoWhat model? And what hardware do you run it on? I find these style of models are great, but fail hard, and fail randomly. I'd be hesitant to use it for a daily driver, but I'm using dual 3060s, so it's not like I'm quantizing a frontier model here. How do you find the overall experience? And do you have any special sauce or recommendations for going this route?
- kamranjon 3mo agoI’m using DeepSeek V4 Flash on 128gb mbp - it’s a bit different using a 200b+ param model. It’s MoE so performance is acceptable. It will still malform a tool call every now and then, but the capabilities are so far ahead anything else that the majority of the time it works really well and solves really complex problems.
- apitman 3mo agoI've got dual 3060s as well. What's the best models you've found for this setup?
- RALaBarge 3mo agoWhat have you been using?
- kamranjon 3mo agoDeepSeek V4 Flash with DwarfStar: https://github.com/antirez/ds4 https://github.com/antirez/ds4 The 2 bit quants are really good. I have a lot of memory so I can squeeze it all in at ~80gb.
- 3mo ago
- SwellJoe 3mo agoWhat I most want to see it compared to is Gemma 4 12B in the 4-bit QAT version. It's barely bigger than this at just under 7GB, so it also runs on just about any modern device and is remarkably smart for its size. It's an excellent tool user, crazy good vision for its size. I'm still trying to wrap my head around how much is lost with each step down in resolution, but the QAT versions from Google seem to prove the answer is "very little" at four bits.
- verdverm 3mo ago4bits is a cutoff point for many model families, but also depends on what parts you quant to 4bits vs alternatives (weights, weight+activation, kv cache). Also depends on model size and task, lots of nuance in quanting I've come to learn. Good evaluation from 2024 https://arxiv.org/pdf/2402.18158 https://arxiv.org/pdf/2402.18158 I'm currently working towards an updated version (not an og author), curious if others are aware of similar surveys, as I have yet to do a real lit search.
- dofm 3mo agoThe key point here, I think, is not the 4-bit but the QAT — the model is trained with the objective of losing the least at 4-bit quantiZation (I am assuming it is literally about assigning numbers that quantize better). The 12B QAT model is indeed sort of mindblowing.
- verdverm 3mo agoI haven't dug into QAT deeply, better recovery is my understanding as well, and also that it is out of reach for most people because you have to train a model to back prop errors based on estimated error under quant. Hopefully more of the lab releases are trained under QAT so we can all benefit.
- dofm 3mo agoI think they did Gemma 3 QAT models and there are QAT versions of essentially all the Gemma 4 models (including DiffusionGemma).
- verdverm 3mo agoPreliminary analysis via lm-evaluation-harness + vllm model | disk | wikitext | gsm8k (match/error) baseline | 55G | 8.00 | 0.50/0.09 nvfp4-gptq | 27G | 8.25 | 0.47/0.9 nvfp4a16-gptq | 27G | 8.11 | 0.53/0.9 bonsai-4bit | 19G | 16.75 | 0/0 (eval bug?) Looks like they quant'd too hard at 4 bits, can't imagine the ternary being any good based on this. I'm also not sure what is up with the gsm8k, their benchmarks show something different, but they are using another eval tool. I'll have to add it to my setup. Also why I'm building a setup instead of taking model devs word for benchmarks. (https://github.com/modelscope/evalscope https://github.com/modelscope/evalscope) Code if you'd like to reproduce or try other test sets: https://github.com/verdverm/quantr https://github.com/verdverm/quantr (lightly tuned to a single oem spark, probably possible in 32-48G) Good paper to understand the effects of quant regimes across model families and tasks: https://arxiv.org/abs/2402.18158 https://arxiv.org/abs/2402.18158 (Evaluating Quantized Large Language Models - 2024 ICML)
- kamranjon 3mo agoDoesn’t this suggest you aren’t properly running the model?
- verdverm 3mo agoIts unclear if its in vllm, the eval harness, or the weights. I'm running it the same as other qwen3.6 derived models and it appears to work, but other comments speak of waiting for support in their tool. Could be a chat template or could be legit because I'm using a different quant / bit selection from their hugging face and using their primary model as the comparison point for my huh? They may have put more effort into the flagship model than the 4bit because they are focused on on-device model running. Software was already complex, now we are adding highly non deterministic elements into the mix
- drob518 3mo agoThis is going in a good direction.
- Luker88 3mo agoNice! Do they have plans to bring even bigger models down to ~16GB VRAM so that more consumer hardware might be useful?
- verdverm 3mo agobigger quant'd harder is not always better than a model of more modest size and quant
- trvz 3mo agoI still don’t see the point of this. In my testing, it’s worse than Qwen 3.5 4B and even 0.8B.
- rvz 3mo agoIt is still worth experimenting. Although we do need independent evaluation of these models instead of labs posting such biased and skewed results.
- kamranjon 3mo agoWhen new models are released (I realized this is qwen 3.6 but the quant is novel) - it takes a few days for the kinks to get worked out - you’ll likely have better luck if you give it a few days and try again.
- athrowaway3z 3mo agoSo first off, phenomenal stuff to see a 1bit model at 90% capability. However, this is the 5th product post in 2 weeks that proclaims that AI use is shifting, and why [insert tradeoffs] are the perfect fit. Paradigms shift don't happen in the release announcements. I suspect this is an AI-ism, making all the release posts sound so paradigmshiftery.
- jrflo 3mo agoThat is true of all tech announcements. Marketers will do their thing regardless of reality.
- nttylock 3mo ago[flagged]
- bilsbie 3mo agoWhat data type stores one bit? Does it offer opportunity for more efficient matmuls?
- Heliodex 3mo agoNice to see a larger model in their lineup, I've been using Ternary 8B and it seems to get higher TPS than most other similarly sized models on my hardware.
- RugnirViking 3mo agomaybe its nitpicking here but the demo shows them asking the model what to cook and its recipie sounds like it wouldn't be very good and also it totally gets the macronutrients wrong. 25g protein for "spaghetti, carrots, peppers, garlic and herbs"?
- nozzlegear 3mo agoI personally don't like carrots much, but it doesn't sound bad to me – could definitely use a tomato sauce but that doesn't seem to be an option from the image it was given. > 25g protein for "spaghetti, carrots, peppers, garlic and herbs"? Maybe it assumed the pasta was some kind of protein chickpea pasta? =P it definitely seems wrong.
- SamBam 3mo agoMany people don't realize that regular spaghetti actually is pretty high in protein. All that gluten.
- sandworm101 3mo agoAnd why would i want such mundane questions to be handled by an AI on my phone? That sort of thing doesnt need AI, let alone a local one. Basic google search was answering those questions long ago. My point: phone-sized AI is only useful if it can do things that only AI can do. Can it ingest a document scanned by the phones camera? Can it translate in real time? I dont see how or why i would ever ask it for recipe advice. That need is met elsewhere x10.
- recursivegirth 3mo agoAwwwe, cmon. You are thinking about the problem space incorrectly. This is an opportunity to create a unicorn company that develops the first AI tongue.
- codebje 3mo agoIf it can give me the recipe without 14 pages of backstory about how Nonna used to make it, it'd be satisfying a real need.
- armanj 3mo agoBonsai vs Qwen (quick) Benchmark: https://github.com/ArmanJR/PrismML-Bonsai-vs-Qwen3.5-Benchmark https://github.com/ArmanJR/PrismML-Bonsai-vs-Qwen3.5-Benchma...
- verdverm 3mo agoBonsai is qwen3.6 based, not 3.5 Likely apples / oranges
- armanj 3mo agoBonsai 8B and 1.7B were on Qwen3.5 the benchmark is from a few months ago. However I'll add Qwen3.6 to the benchmark too.
- nostrebored 3mo agoqwen3.6 starts at 27B
- raylad 3mo agoNot impressed. It fails the "Jabberwocky" test.
- raylad 3mo agoThis got a downvote and I understand why: because I didn't describe the test, which is to ask it "Please recite Jabberwocky". This is actually difficult because there are so many invented words in the poem which have extremely low frequencies in the training data. So a model that can do it properly is likely to be very good in other ways. Qwen-3.6-27B can do this until it gets overly quantized.
- est 3mo agoMaybe Taalas could cook this as their AI-on-chip next
- diddid 3mo agoThey should do this to GLM 5.2
- KolinFirz 3mo ago[dead]
- luciana1u 3mo ago[flagged]
- all2 3mo agoTried this on an old 4 core i5 and got about 1tps. OS: WSL2 on Windows 10
- ExxKA 3mo agoFrom an investors perspective, this is truly a paradigm shift - this will kill a whole range of startups in Europe which were packaging privacy and wrapping around large hosted models. There's absolutely no reason to use a "Privacy GPT tm" provider, then I have it all on my own laptop - There is also no need for banks or other regulated institutions to rely on those providers when they can selfhost with this much intelligence on tap.
- dofm 3mo agoIt will when they can get the performance up a bit. My brief experiments with the ternary version suggest that it broadly meets their claim to be a 27B model that fits in much less RAM, that is for sure. It is about as fast as the underlying Qwen 27B but it gets stuck in reasoning loops quite easily.
- unitindex 3mo ago[flagged]
- yieldcrv 3mo agothis is really amazing! keep pushing guys, this will coexist in the memory footprint with vision models, and audio models, and other kinds of transformers so we still need memory to work with you also might single handedly pop the hyperscaler investment and capital projects! that's the whole AI bubble essentially!
- goofy_lemur 3mo agoI tried this on M1 Pro today with 16GB ram and it worked!!! I was using vscode and it seemed to interperet the system prompt right and then started actually inspecting and doing stuff. Unfortunately the vscode system prompt is 24000 tokens, and I was getting 100 at beginning, 69 by the end of it, but honestly I'm super impressed. Great work team 1
- RandyOrion 3mo agoThanks Bonsai team. Now open weight LLMs/VLMs/LMMs are becoming even larger to the extent that consumer-grade hardware are no longer able to run these models. In contrast, quantization and pruning make the model better at the size-performance pareto and provide people with strictly more possibilities.
- sahitya_ 3mo ago[flagged]
- hham 3mo agoThis is accelerant #3 and #4 from our article converging in one release: a 27B-class model, built on Qwen (already one of our examples of local models "good enough to matter"), now running on an iPhone. The hardware layer and the local-model layer aren't just going to converge in the future, they're doing it right now! https://news.ycombinator.com/item?id=48892559 https://news.ycombinator.com/item?id=48892559
- ryss20 3mo ago[flagged]
- deleted 3mo ago[deleted]
- runtime_lens 3mo ago[flagged]
- lifesucks1 3mo agoWhen Bonsai GLM 5.2 2bit
- networked 3mo agoI have benchmarked Bonsai 27B CPU inference on my computer (a Ryzen 7 5700X desktop with 48G RAM running Ubuntu 24.04) using the latest 62061f910 build of PrismML's llama.cpp fork. Binary: 9 t/s prompt, 6 t/s generation. Ternary: 0.8 t/s prompt, 0.7 t/s generation. It looks like CPU inference for ternary isn't optimized yet.
- rvba 3mo agoDoes anyone know how to disable the 10 minute timeout in Android studio when using a self hosted model on local machine? It always auto disconnects.
- anshumankmr 3mo agoMore and more it seems the iPhone 16 was the worst deal in history cause I don't think mine will support the upcoming foundation model from Apple or this one, does it?
- davedx 3mo agoI find it super interesting that we're now in an era where we have LLM's that are quantized to binary weights - 1's and 0's. So effectively they're digital neural networks. I assume that in addition to the significant memory savings, this should also lead to much simpler matrix multiplication operations? Could models like these run on CPU's efficiently, or does the geometry of the compute mean GPU's are still a better choice?
- jboss10 3mo agoMost of the time, the speed of these models are constrained by memory bandwidth. GPUs normally have much more memory bandwidth.
- davedx 3mo agoI'd expect the memory bandwidth to be the same for the CPU and GPU under a unified memory architecture like Apple silicon uses?
- kmacdough 3mo agoIts not about geometry, it's a parallel compute thing thing. CPUs typically don't have more than 10 or 20 cores. GPU have 100s to 1000s. Matmul is very well parallelized. More, lower power cores will always pay off handsomely.
- snthpy 3mo agoWhat is the best way to deploy these on CPUs, arm64 ones in particular? I'm interested in the CPU inference application of these models with things like the FairyFuse kernels. I've tried trillim previously but was disappointed that i got higher tok/s just with similar sized models through ollama using just Q4_K_M quants. I see there is bitnet.cpp and litespark-inference. What else should i look at?
- MrDrMcCoy 3mo agoLlama.cpp should be rolling out support for this soon if they haven't already. Cactus is a but more targeted for efficient ARM execution, but I haven't been keeping up with what all they support. Would be worth a feature request if they aren't working on it yet.
- OsamaJaber 3mo ago[dead]
- latexr 3mo agoThe meal demo is hilarious. — Hey, model, see this fake-ass stock photo of a variety of spices, vegetables, and spaghetti? What meal can I make with this? — Just cook everything. — I’m a complete noob. I can’t even fathom how to cook those things. Help me! — Sure sure. First boil the spaghetti completely and drain. Only after that, while it’s getting cold, you need to sauté (good luck knowing what that is if you don’t even know how to cook spaghetti) the garlic and carrots at this specific temperature (good luck figuring out how to do that on a stove). Despite having mentioned the peppers and herbs in the previous message, I’m not going to tell you what to do with those. Just chew them raw or something, I guess. The demo shows that the model can answer, but the answers are frankly bad. Here’s what you could’ve done instead faster with better results: a web search for “spaghetti carrots peppers”. Don’t even need to add “recipe”. Presumably you’ve been using the model as you develop it, why not show something real and useful instead of a generic, unrealistic and uninteresting scenario that above all makes it look incompetent? Show something that genuinely surprised you positively.
- kbart 3mo agoExcuse my likely stupid question, but has anybody had some success using Claude Code with frontier agents (or Junie or anything else) to invoke local LLMs for specific sub-tasks or wrapped as skills? In other words, is there a way to use expensive, frontier models as orchestrators that manage local models to do the specialised coding tasks?
- adam_patarino 3mo agoWhat for? Cost savings? Something else?
- docheinestages 3mo agoLook into FastContext by Microsoft. Not extraordinary but specifically designed for paired usage with a larger LLM to save tokens [1]. [1]: https://github.com/microsoft/fastcontext https://github.com/microsoft/fastcontext
- HnUser12 3mo agoThat repo suddenly seems to have gone missing. I get 404.
- docheinestages 3mo agoOh, I just realized Microsoft removed it a couple weeks ago. I had the link in my bookmarks. The model is available on HuggingFace [1]. [1]: https://huggingface.co/models?sort=trending&search=fastcontext https://huggingface.co/models?sort=trending&search=fastconte...
- rtcoms 3mo agoThis may be useful in this context: https://entelligentsia.github.io/is-grep-enough/fastcontext.html https://entelligentsia.github.io/is-grep-enough/fastcontext....
- fabioz 3mo agoI'm working on the area in https://beolis.com https://beolis.com. The system as a whole is meant to support that use case, where each task (ticket in its jargon) can be tackled using a custom workflow that can each use a different agent/llm (so, it should support local LLMs if you have configured your coding agent to use them). Sidenote: it's still not where I want, but getting there...
- Havoc 3mo agoGot this running on my phone. Unfortunately like other small models it hallucinates quite easily. eg asked it what Signoz is. It reckoned it is a woocommerce/shopify competitor aimed at India market
- Havoc 3mo agoSame for Talos. Response is 100% hallucinations
- contentpulse 3mo ago[flagged]
- stevenhubertron 3mo ago[flagged]
- OutOfHere 3mo agoHow does one even install and use this on a phone? They don't say. I guess the claim is fake. It is unreasonable to make a claim without an app for Android and iPhone that supports and runs this model.