4 ms·
If you want to try out out the GGUFs from https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#these-files-need-our-llamacpp-build https://huggingface.co/p
by simonw 17d ago
If you want to try out out the GGUFs from https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#these-files-need-our-llamacpp-build https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#th... be aware that you need Prism's llama.cpp fork to get them to work, from https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15 https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-...
This should work:
cd /tmp
# Get the Prism macOS runtime
curl -fL https://github.com/PrismML-Eng/llama.cpp/releases/download/prism-b10685-7dffb15/llama-prism-b10685-7dffb15-bin-macos-arm64.tar.gz -o bonsai-runtime.tar.gz
tar -xzf bonsai-runtime.tar.gz
# Get the ~5.95 GB GGUF model:
curl -fL https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/main/Ternary-Bonsai-2-27B-PTQ1_0.gguf -o Ternary-Bonsai-2-27B-PTQ1_0.gguf
# Run the server, I used port 8331
./llama-prism-b10685-7dffb15/llama-server \
-m Ternary-Bonsai-2-27B-PTQ1_0.gguf \
--port 8331 -ngl 99 -fa on -c 32768
Then open http://localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this:
uvx llm openai endpoint http://127.0.0.1:8331/v1 \
--model bonsai-2-27b --responses hi
That's running at ~20 token/second for me on an M5 Pro (after a server restart I got 44 token/second, not sure why), but I'm pretty sure something isn't working right, on startup the server said "ggml_metal_device_init: - the tensor API is not supported in this environment - disabling".
- deleted 17d ago[deleted]
- simonw 17d agoI used that to Generate an SVG of a pelican riding a bicycle: https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fba4cf3a88f4e7dc32994f2672150f770 https://tools.simonwillison.net/markdown-svg-renderer?url=ht... It took 18 minutes 20 seconds. Pretty decent for a 5.5GB model file.
- kadoban 17d agoHonestly looks pretty good except whatever is going on with its booty. Is that an ass helmet? I cannot parse what's going on there.
- Forgeties79 17d agoI think it’s supposed to be a wing
- bigwheels 17d agoI like the lens effect behind the rear tire.
- tomcam 17d agoLike you've never worn an ass helmet
- kadoban 17d agoOnly because I hadn't previously thought of it xD Step up from the standard ass-hat for sure.
- deleted 17d ago[deleted]
- rahimnathwani 17d agoM1 Pro, same prompt, same cli options: 32,706 tokens 38min 19s 14.22 t/s
- raylad 17d agoHow does that compare with the bf16 version? For my "Please recite Jabberwocky" test the bf16 almost passes but the ternary and even fp8 versions fail badly.
- ctolsen 16d agoDefinitely a little better. https://gist.github.com/ctolsen/b2883e7cbf5e4357fa04366019e60bfe https://gist.github.com/ctolsen/b2883e7cbf5e4357fa04366019e6...
- shmoil 16d agoCan you ask it for an SVG of a bicycle riding a pelican? Thanks.
- refibrillator 17d agoWhere did you get these instructions? They have a demo repo with a setup.sh script: https://github.com/PrismML-Eng/Bonsai-demo https://github.com/PrismML-Eng/Bonsai-demo The release tag and weight file you suggest doesn’t match what they wrote.
- simonw 17d agoI figured them out, starting from the GGUF on Hugging Face. If you have found better instructions and they work then use those instead! Personally I prefer to download models directly rather than running some `./setup.sh` script where I need to then review what it does first.
- refibrillator 17d agoYeah just wanted to mention in case it explains the 2x lower throughout you are seeing on M5. To be fair their documentation is a bit inconsistent in some spots. Would be good to know if the release and weights from their demo repo work better. I’m trying on a 4090 and will report back.
- fnordpiglet 16d agoThey’re fairly useful as constrain domain classifiers due to the low memory requirements means you can stuff a lot of them into a less expensive GPU farm and get really decent throughout with pretty good results over all. At least that’s my experience. I wouldn’t bother using a tiny model for coding - but the world is full of abductive reasoning tasks that don’t involve coding.
- nikwen 17d agoIt would be great to have upstream llama.cpp support for this!
- rahimnathwani 17d agoIf you want to download the gguf to your regular huggingface cache directory instead of to /tmp, you can download the model and run the server in one step: export HF_TOKEN=xxx # optional, speeds up the download ./llama-prism-b10685-7dffb15/llama serve \ -hf prism-ml/Ternary-Bonsai-2-27B-gguf:PTQ1_0 \ --port 8331 -ngl 99 -fa on -c 32768
- francisjp 17d agoThanks for all of your exploration in public Simon. Commenting because the fix I proposed was merged in roughly 49 commits after the PrismML Fork. The “tensor API is not supported” warning occurred because llama.cpp’s startup probe fails to compile a matmul2d kernel: Metal’s tensor headers require language version 4.0, but ggml-metal-device.m previously omitted MTLCompileOptions.languageVersion, disabling the API universally. Here’s a link to the diff if you want to try and update that fork to take advantage of the prefill gains afforded by the hardware: https://github.com/ggml-org/llama.cpp/pull/27461/changes https://github.com/ggml-org/llama.cpp/pull/27461/changes
- jb_briant 17d agoThat kind of issue is exactly why Im so happy to have LLMs, let it take one hour or trial and error instead of me spending a day digging traces
- wombat23 17d agoI managed to run it with RTX 3070 (8GB VRAM) following the "Quickstart" on HF model card with minor modifications (modify the architecture 86 for your own hardware): git clone https://github.com/PrismML-Eng/llama.cpp && cd llama.cpp cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86 && cmake --build build -j Then downloaded & verified `Ternary-Bonsai-2-27B-PTQ1_0.gguf` from HF and ran ./llama.cpp/build/bin/llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -fa on -c 32768 -b 256 -ub 64 --temp 1.0 --top-p 0.95 --top-k 20 -ctk q8_0 -ctv q8_0 -p "hello world in x86 assembler" -n 256 the parameters were suggested by gpt-5.6-luna to reduce memory footprint, as the defaults ran OOM on my gpu. result looks good: [ Prompt: 165.6 t/s | Generation: 40.6 t/s ] would be nice if they upstreamed their changes so that it runs with the original llama.cpp
- aktenlage 17d agoWould it speed up prompt processing if you increased the -ub (and -b) parameters.
- wombat23 16d agoI don't see any significant speed up. also with defaults according to --help -b, --batch-size N logical maximum batch size (default: 2048) -ub, --ubatch-size N physical maximum batch size (default: 512) I also tried double the default. that also means that they can be left to default settings, apparently.
- ekianjo 16d agowow why is prompt processing so slow?
- zepearl 16d agoExact same test executed on my RTX 3060 (12 GiB VRAM, PCIe 3.0 4x slot): [ Prompt: 95.0 t/s | Generation: 26.5 t/s ] (the test's prompt is very short but with longer ones the I get ~200 prompt processing rate, but I was hoping for a better token generation rate...) Am I understanding correctly that no draft model exists (will never exist or just currently does not exist yet)? There is no draft file in Huggingface's repository ( https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/tree/main https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/tr... ) and in the file "scripts/download_models.sh" of the demo repository ( https://github.com/PrismML-Eng/Bonsai-demo/blob/main/scripts/download_models.sh https://github.com/PrismML-Eng/Bonsai-demo/blob/main/scripts... ) I see this remark: if [ "$_family" = "bonsai2" ]; then the projector ships in the same repo; Bonsai 2 has no dspark drafter
- ricardobayes 16d agoThanks for this, I never knew llama-server has a web ui until now.
- jakswa 16d agoI love their web UI so much that I had AI slop all over it in a fit of fanboy-ism https://inkcap.click https://inkcap.click
- ithkai92 16d agoThanks for the headstart, I saw hf also has PQ2_0 and able to finetune the command and in a MBA M4 24GB averages around 10t/s with the command. ./llama-prism-b10685-7dffb15/llama-server \ -m Ternary-Bonsai-2-27B-PQ2_0.gguf \ --port 8331 -ngl 99 -fa on -c 65536 --jinja \ --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0