3 ms·
Exact same test executed on my RTX 3060 (12 GiB VRAM, PCIe 3.0 4x slot): [ Prompt: 95.0 t/s | Generation: 26.5 t/s ] (the test's prompt is very short but wi
by zepearl 13d ago
Exact same test executed on my RTX 3060 (12 GiB VRAM, PCIe 3.0 4x slot):
[ Prompt: 95.0 t/s | Generation: 26.5 t/s ]
(the test's prompt is very short but with longer ones the I get ~200 prompt processing rate, but I was hoping for a better token generation rate...)
Am I understanding correctly that no draft model exists (will never exist or just currently does not exist yet)?
There is no draft file in Huggingface's repository ( https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/tree/main https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/tr... ) and in the file "scripts/download_models.sh" of the demo repository ( https://github.com/PrismML-Eng/Bonsai-demo/blob/main/scripts/download_models.sh https://github.com/PrismML-Eng/Bonsai-demo/blob/main/scripts... ) I see this remark:
if [ "$_family" = "bonsai2" ]; then
the projector ships in the same repo; Bonsai 2 has no dspark drafter