3 ms·
I used devstral today with cline and open hands. Worked great in both. About 1 minute initial prompt processing time on an m4 max Using LM studio because the
by zackify 1y ago
I used devstral today with cline and open hands. Worked great in both.
About 1 minute initial prompt processing time on an m4 max
Using LM studio because the ollama api breaks if you set the context to 128k.
- elAhmo 1y agoHow is it great that it takes 1 minute for initial prompt processing?
- cheema33 1y agoThat time is just for the very first prompt. It is basically the startup time for the model. Once it is loaded, it is much much faster in responding to your queries. Depending on your hardware of course.
- zackify 1y agoHaha great as in surprisingly good at some simple things that nothing has been able to do locally for me. The 1 minute first token sucks and has me dreaming for the day of 3-4x the bandwidth
- nico 1y agoHave you tried using mlx or Simon Wilson’s llm? https://llm.datasette.io/en/stable/ https://llm.datasette.io/en/stable/ https://simonwillison.net/tags/llm/ https://simonwillison.net/tags/llm/
- zackify 1y agoOn lm studio I was using mlx