6 ms·
1-Bit LLM in the Browser
- DerNico 3mo agoDoes not work Error: Unsupported device: "webgpu". Should be one of: wasm.
- stfurkan 3mo agoI'm currently working on an open-source engine [1] exactly for this purpose. If anyone wants to try or has any suggestions, I'm happy to listen :) 1. https://github.com/stfurkan/bitgpu https://github.com/stfurkan/bitgpu 2. https://aidekin.com https://aidekin.com --> this is one of my projects that's currently using bitgpu engine
- mgoetzke 3mo agoInteresting. But trying the demo with 27G model and simple 8+5 did not work. Guessing there is some work to do still ``` calculate({"expression":"g"}) → error: only numbers, + - * / and parentheses are allowed error: No user query found in messages. What is 8+5? Use the calculator get_time({}) → Mon Jul 20 2026 10:36:28 GMT+0200 (Central European Summer Time) error: No user query found in messages. hi Hello! 2 tokens · 6.2 tok/s · prefill 4563 ms · confidence 87% which tools do you know about ? I'myepic know about the following are two tools I am aware of course, I currently i know about I know about me know I am aware of course I know I know I am aware of the following calculate({"expression":" I know I am aware of know I am aware of I am aware I know I know know "}) → error: only numbers, + - * / and parentheses are allowed error: No user query found in messages.```
- stfurkan 3mo agoThank you for trying it :) I just added the 27B support 1 hour ago. Let me check and try to fix it. All other models should work as expected but if you have issues with them too, I can take a look.
- stfurkan 3mo agoHi again :) I updated the engine and 27B model should work in the demo[1]. We can still consider it as experimental but I am trying to make it better and efficient. 1. https://stfurkan.github.io/bitgpu/examples/chat.html https://stfurkan.github.io/bitgpu/examples/chat.html
- alex7o 3mo agoHay man, I really want a bonsai style embeddings or re-ranking model for you think sth like that is feasible?
- stfurkan 3mo agoHi :) There are some smaller embedding and re-ranking models (I am currently using some of them on my aidekin project) but having it like bonsai models will be great to have.
- luca-3 3mo agoThis is great for high-privacy use cases. If you could add an agentic loop that can call tools within the website using the user’s session, that would be great. Then the user can chat with an AI about sensitive information in the most secure way possible. Also, do you have data on the technical requirements for this? If someone uses an old phone, e.g., does it still work?
- stfurkan 3mo agoHi, the engine itself already has tool calling capabilities but if you're asking for the aidekin I didn't add tool calling there on purpose. It needs webgpu support but I haven't had a chance to test with different devices.
- EGreg 3mo agoYes, I’m very interested in exploring this stuff. Want to connect and discuss it? https://calendly.com/safebotsai https://calendly.com/safebotsai
- stfurkan 3mo agoHello, I sent you a LinkedIn connection request. We can chat first and schedule a meeting. Thanks
- espetro 3mo agoCool stuff! I see you're targeting two verticals here (runtime on top of ONNX + TS library), I'm curious did you consider simply creating a AI SDK Provider (https://ai-sdk.dev/providers/community-providers/custom-providers https://ai-sdk.dev/providers/community-providers/custom-prov...) instead of crafting a separate library? Besides it, how are you dealing with the harness the model can use on web? I'm currently building https://github.com/bolojs/bolo https://github.com/bolojs/bolo to extend the base harness agents can use on web and I'm interested in hearing from your experience on crafting tools for bitgpu
- stfurkan 3mo agoThank you :) I wanted to create it from zero myself so I can control it better. e.g. I started with ONNX support but now it support gguf models too. It might be a good idea to add some popular connectors in the future. Good luck with your project :)
- lnenad 3mo agoCool demo, but even the smallest model does 2.5 tps on my 7900xtx which is not great.
- k3liutZu 3mo agoI get 9-10 tps on my MacBook Pro. Though the small model seem really limited
- whstl 3mo agoInterestingly, I get 44 tps on Safari, MacBook Pro base model.
- Tade0 3mo ago60tps in Chrome here - 14" MBP M4. Safari loads 3% of the model and then stops - repeatedly. Basically zero using a 7700S in Chrome on Ubuntu though. Must be not working at all, as the performance difference can't be that large.
- christkv 3mo agoHmm I get 44,7 tokens/s on the smallest 1.7b one in the browser on my m1 max.
- LorenDB 3mo ago18.8 tps on an 11th gen mobile i7 here.
- dim13 3mo agoFails "car wash" test miserably. Ask it: > I want to wash my car. The car wash is 50 meters away. Should I walk or drive? Ref: https://opper.ai/blog/car-wash-test https://opper.ai/blog/car-wash-test
- HelloUsername 3mo ago[dead]
- SwellJoe 3mo agoEven very large models failed tests like this up until a year or two ago...this is a very, very, small model.
- woadwarrior01 3mo agoWorks with the Bonsai-27B-1bit-mlx model on my mac. Edit: this is running with MLX, and not with WebGPU on the browser, also it is a fail. """ You should *walk*. Here’s why: - *Distance*: 50 meters is very short — about 0.3 miles or 0.5 kilometers. - *Time*: Walking takes roughly 2–3 minutes. Driving would take longer due to parking, starting the engine, and maneuvering. - *Effort*: Walking is light exercise and avoids parking hassle. - *Safety*: Less risk of misjudging parking space or damaging the car. Unless you have a disability that makes walking difficult, walking is the most practical and efficient choice. """
- Xalutiono 3mo ago
- indianmouse 3mo agoCould not get any usable output from it for any real usecase. Can't see any real value in using 1 bit llms except and apart for the sake of running it for demo or feel good factor of running a higher param model locally and in limited hardware! Since use as it is very limited, could someone throw some light on any actual use case?
- richrichardsson 3mo ago[dead]
- deleted 3mo ago[deleted]
- Lwerewolf 3mo agoPretty sure this might be a duplicate. Regardless, tried the 1bit bonsai 27b gguf three different ways - their llama.cpp fork (prism, was it) on 2 machines (1255u/16g, m5 max/128g) and the web (on the 1255u). With llama.cpp it works. on the web it emits the same thing over and over. Prompt was "do you see anything wrong with this code: <C code, compound literal passed as a pointer to a function in an if statement>", result was literally "Yes, let'ssomestruct_struct_tsomestruct<a bit of similar garbage>structstructstruct<forever>. Locally, it's definitely not the full 3.6 27b, but for ~6G with the context, it's quite impressive, and I like what this implies for larger models. Speeds on the 1255u are abysmal, though (~1TPS or less) - granted, CPU-only.
- kees99 3mo agoQwen is quite slow in general. I'm getting ~0.5 tps from [qwen], and ~10 tps from [gemma] - roughly the same size, same quant, same hardware (8-core CPU) same software (llama.cpp). [qwen] Qwen3.6-27B-Q4_K_M.gguf [gemma] gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf
- christkv 3mo agoQuite different models though Qwen 3.6 model has 27B active parameters. gemma-4 has 26B parameters with 4B active at any given time. You should compare with Qwen 3.6 35B A3B
- deleted 3mo ago[deleted]
- einpoklum 3mo agoI got 1.
- estebarb 3mo agoI tried the smallest ome, the only available, and got an error: ``` Could not load. Error: failed to call OrtRun(). ERROR_CODE: 1, ERROR_MESSAGE: Non-zero status code returned while running GroupQueryAttention node. Name:'/model/layers.0/attn/GroupQueryAttention' Status Message: Failed to create a WebGPU compute pipeline: A valid external Instance reference no longer exists. ```
- leumon 3mo agoyou need to enable webgpu
- PeterSmit 3mo agoCool idea. Crashes on iOS though, after downloading the 1.7B model
- dwa3592 3mo ago[dead]
- lelanthran 3mo ago1-bit is not much though. Here's what I got: Me: > Describe the process of pasteurisation Response: > pasteurization is a, which a, the which which is, the and process, the the past, the and the, and and and past, and past, and and and and the, the and past, process is the is is is is, and the, past, and and future, process, and and and and and and and the, past, process, process, the the, the the, is and the, the the, and and and, and process, and and the, past, and and and past, and the, the, and the past, the the, process, process, and past, the past, past, the and, the past, and and and and and and and and and the, and and and and and and and, the the, the the, the or and and, the the, the and the, which the, and the, the the, past, and n the process, and and and and, past, and, and and, the past, and, the the, past,, and the, the the, the is the, past, and and and and and, and and and past, the and the, the the, the the, are and past, and which the, the and n, n the, the n past, past, n the, and n, the the process, which past, the the, the n, the the, the is past, the the, is, past, the the, past, and past, process, the the, the the, the and and the, and which past, the (and that basically just goes on and on like that)
- wongarsu 3mo agoI got a decent response with the same prompt But yes, this will happen. 1.7B is already a tiny model, running that at 1-bit is going to lead to some behavioral issues and will need tuning of temperature/topP/repeatPenalty/etc, and probably a software fallback that detects pathological cases like this and simply retries. Just like we all had to do in the times of GPT 2 and GPT 3
- OutOfHere 3mo agoI confirm I got an extremely detailed and long response for "Describe the process of pasteurization." with 1.7B. I had to tune nothing.
- nikhizzle 3mo agoIMHO from guy working with heavily quantized models and trying to get results. I do think quantization at less than 4-bit requires more advanced mathematics than is usually used. See Google Scholar Babak Hassibi (founder of Prism ML). Generally am amazed about the resiliency of these new data structures to approximation.
- leumon 3mo agoi found this space to work better/at all: https://huggingface.co/spaces/webml-community/bonsai-webgpu-kernels https://huggingface.co/spaces/webml-community/bonsai-webgpu-...
- spuz 3mo agoUnfortunately, this doesn't work for me. After loading for the first time, my first prompt had it generate an infinite series of exclamation marks. My subsequent queries just had it return nothing. I have plenty of RAM and VRAM. This is with Chrome on Linux Mint with 32GB RAM and a 12GB 3060.
- om8 3mo agoThis project needs webgpu -- I did it on cpu about a year ago. My demo uses 2 bit quantization to run llama3 models on any device with enough ram. https://galqiwi.github.io/aqlm-rs/ https://galqiwi.github.io/aqlm-rs/
- aavci 3mo agoOther than privacy and security and simplified development, what are other valid reasons you'd prefer to run this instead of a backend LLM?
- OutOfHere 3mo agoI run Bonsai-27B Q1_0 on my Android phone via the PocketPal app. It's slow but good. I wouldn't want to run fewer B. By any measure, my phone has less memory than my laptop browser.
- boramdd 3mo ago[flagged]
- totallygeeky 3mo agoImpressive, I asked it how many r's in strawberry and it was immediately wrong with 2. any further questions I asked it proceeded to re-explain why there were 2 r's in strawberry, despite asking the silly "walk or drive to get my car washed" question. I struggle to see what these sorts of models could possibly be good at.
- le_master_code 3mo ago[dead]