3 ms·
Wow, amazing! What if there is enough RAM to fully load the model? I assume in that case I shouldn’t use your engine.
by yakupov_bulat 2mo ago
Wow, amazing!
What if there is enough RAM to fully load the model? I assume in that case I shouldn’t use your engine.
- 0gs 2mo agoyou could use mine ... github.com/0gsd/enough (it has other stuff too)
- gitpusher42 2mo agoIt depends on the use case. I measured this exact model with a 4k context on the mlx engine. It runs at 75 tok/s on my M5 Mac Pro and using 14 GB of RAM. For my engine the same model uses 2 GB of RAM and produces 31–35 tok/s. The project is still experimental so performance may vary as it continues to improve. If you want to save around 12 GB of RAM for other tasks and you are ok with 35 tok/s (afaik it is roughly comparable to ChatGPT’s speed for basic responses) my engine may be a good fit. If you need maximum speed and flexibility just use MLX
- anentropic 2mo agocan I vary the context length depending on RAM available?
- gitpusher42 2mo agoYeah, sure! You can select different options in the app settings at the right panel, it shows how much memory it will use For CLI and Server, use --max-context