3 ms·
I managed to run it with RTX 3070 (8GB VRAM) following the "Quickstart" on HF model card with minor modifications (modify the architecture 86 for your own hardw
by wombat23 14d ago
I managed to run it with RTX 3070 (8GB VRAM) following the "Quickstart" on HF model card with minor modifications (modify the architecture 86 for your own hardware):
git clone https://github.com/PrismML-Eng/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86 && cmake --build build -j
Then downloaded & verified `Ternary-Bonsai-2-27B-PTQ1_0.gguf` from HF and ran
./llama.cpp/build/bin/llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -fa on -c 32768 -b 256 -ub 64 --temp 1.0 --top-p 0.95 --top-k 20 -ctk q8_0 -ctv q8_0 -p "hello world in x86 assembler" -n 256
the parameters were suggested by gpt-5.6-luna to reduce memory footprint, as the defaults ran OOM on my gpu. result looks good:
[ Prompt: 165.6 t/s | Generation: 40.6 t/s ]
would be nice if they upstreamed their changes so that it runs with the original llama.cpp
- aktenlage 14d agoWould it speed up prompt processing if you increased the -ub (and -b) parameters.
- wombat23 14d agoI don't see any significant speed up. also with defaults according to --help -b, --batch-size N logical maximum batch size (default: 2048) -ub, --ubatch-size N physical maximum batch size (default: 512) I also tried double the default. that also means that they can be left to default settings, apparently.
- ekianjo 14d agowow why is prompt processing so slow?
- zepearl 14d agoExact same test executed on my RTX 3060 (12 GiB VRAM, PCIe 3.0 4x slot): [ Prompt: 95.0 t/s | Generation: 26.5 t/s ] (the test's prompt is very short but with longer ones the I get ~200 prompt processing rate, but I was hoping for a better token generation rate...) Am I understanding correctly that no draft model exists (will never exist or just currently does not exist yet)? There is no draft file in Huggingface's repository ( https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/tree/main https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/tr... ) and in the file "scripts/download_models.sh" of the demo repository ( https://github.com/PrismML-Eng/Bonsai-demo/blob/main/scripts/download_models.sh https://github.com/PrismML-Eng/Bonsai-demo/blob/main/scripts... ) I see this remark: if [ "$_family" = "bonsai2" ]; then the projector ships in the same repo; Bonsai 2 has no dspark drafter
- jimmySixDOF 14d agohummm wonder what the context length limit will be like this
- wombat23 14d agoUPDATE: I did more experiments - this time with llama-server and pi harness. i'll leave the results here for posteriority: ./llama.cpp/build/bin/llama-server \ -c "$context_size" \ -ctk q4_0 \ -ctv q4_0 \ --no-models-autoload \ --models-dir ~/bonsai \ --reasoning-preserve \ -ngl 99 the max context size i could serve is 64K on GPU only (the -ngl 99 setting). tested with pi harness and it is very fast. the /thinking level always gets reset to off though and it is not very smart like this. haven't figure out a way to fix that.