Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
shironnnn_
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
shironnnn_
4mo ago
if on MacOS I recommend llm-mlx which currently renders tokens 10%-15% faster than llama.cpp.
2.
▲
by
shironnnn_
4mo ago
I use SpecKit to create a very detailed plan with a high amount of specificity using paid Claude plan. Then I give it to local LLM (eg: Qwen / Gemma 4) via CLI. This is possible through usage of llm-mlx on Mac (or ollama on any machin