5 ms·
It's fast enough for me to cancel monthly AI services on a mac mini m4 max.
by bastardoperator 2y ago
It's fast enough for me to cancel monthly AI services on a mac mini m4 max.
- staticman2 2y agoSmaller, dumber models are faster than bigger, slower ones. What model do you find fast enough and smart enough?
- Matl 2y agoNot OP but I am finding the Qwen 2.5 32b distilled with DeepSeek R1 model to be a good speed/smartness ratio on the M4 Pro Mac Mini.
- a1o 2y agoHow much RAM?
- bastardoperator 2y agoI'm running the same exact models.
- diggan 2y agoCould you maybe share a lightweight benchmark where you share the exact model (+ quantization if you're using that) + runtime + used settings and how much tokens/second you're getting? Or just like a log of the entire run with the stats, if you're using something like llama.cpp, LMDesktop or ollama? Also, would be neat if you could say what AI services you were subscribed to, there is a huge difference between paid Claude subscription and the OpenAI Pro subscription for example, both in terms of cost and the quality of responses.
- deleted 2y ago[deleted]
- fetus8 2y agoHow much RAM are you running on?
- lostmsu 2y agoHm, the AI services over 5 years cost half of m4 max minimal configuration which can barely run severely lobotomized LLaMA 70B. And they provide significantly better models.
- jamesy0ung 2y agoI presume you're using the Pro, not the Max. Anyways, what ram config, and what model are you using?