3 ms·
I just used it on a Apple M4 MacBook Pro with 48GB RAM with llama.cpp and Pi to help diagnose an infinite looping request in a React Server component on a Next.
by paulbjensen 2mo ago
I just used it on a Apple M4 MacBook Pro with 48GB RAM with llama.cpp and Pi to help diagnose an infinite looping request in a React Server component on a Next.js application.
After about 10+ hours of digging, it has apparently found a bug in the Next.js framework, with an example app that replicates the bug, and a fix for now to disable prefetch in the Link component.
I had in my prompt asked it to discover the root cause of the bug and propose a fix, but I did not expect it to dig this deep.
- busfahrer 2mo agoI am eyeing one of these specifically for this use case, could you please post roughly what kind of tokens per second numbers you get for text generation for this 27B model? edit: and which quant you are using, please :-)
- noman-land 2mo agoUsing the 4bit quant on an M1 64GB I'm getting ~65 tps for prompt processing and ~11 tps token generation using oMLX to serve the models and pi as a harness.
- digidecode 2mo ago10 hours at what tokens per sec?
- paulbjensen 2mo agoI was using this model: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF https://huggingface.co/unsloth/Qwen3.8-27B-GGUF Quantization is Q4_K_M (4-bit K-quants, medium) From Pi, these are the up/down token counts: Tokens: ↑147k ↓110k R26.5M Prompt submitted at 18:52:54 BST on Sunday 16th August 2026, and finished at 04:05:27 BST on Monday 17th August 2026. Last print out from the llama.cpp server logs: 881.42.815.480 I slot print_timing: id 2 | task 110233 | prompt eval time = 2742435.76 ms / 134911 tokens ( 20.33 ms per token, 49.19 tokens per second) 881.42.815.483 I slot print_timing: id 2 | task 110233 | eval time = 6312248.28 ms / 5804 tokens ( 1087.57 ms per token, 0.92 tokens per second) 881.42.815.483 I slot print_timing: id 2 | task 110233 | total time = 9054684.04 ms / 140715 tokens 881.42.815.484 I slot print_timing: id 2 | task 110233 | graphs reused = 114795 881.42.820.381 I slot release: id 2 | task 110233 | stop processing: n_tokens = 140714, truncated = 0