3 ms·
Thanks for your recommendation! I just ran Llamafile for the first time with a custom prompt on my Windows machine (i5-13600KF, RX6600) and found that it perfor
by HiPHInch 2y ago
Thanks for your recommendation! I just ran Llamafile for the first time with a custom prompt on my Windows machine (i5-13600KF, RX6600) and found that it performed extremely slowly and wasn't as smart as ChatGPT. It doesn't seem suitable for productive writing. Did I do something wrong, or is there a way to improve its writing performance?
- noman-land 2y agoLocal models are definitely not as smart as ChatGPT but you can get pretty close! I'd consider them to be about a year behind in terms of performance compared to hosted models, which is not surprising considering the resource constraints. I've found that you can get faster performance by choosing a smaller model and/or by using a smaller quantization. You can use other models with llamafile as well. They have some prebuilt ones: https://github.com/Mozilla-Ocho/llamafile?tab=readme-ov-file#other-example-llamafiles https://github.com/Mozilla-Ocho/llamafile?tab=readme-ov-file... You can also search for other llamafiles for other models on HuggingFace by using the llamafile tag. https://huggingface.co/models?library=llamafile&sort=trending https://huggingface.co/models?library=llamafile&sort=trendin... And you can download model weights directly and use them by providing an -m flag to llamafile but that's getting a bit less straightforward. https://github.com/Mozilla-Ocho/llamafile?tab=readme-ov-file#using-llamafile-with-external-weights https://github.com/Mozilla-Ocho/llamafile?tab=readme-ov-file...
- jonnycomputer 2y agoRAM and what GPU you have are big determinants of how fast it will run, and how smart a model you can run. A large amount of RAM and GPU memory is required for larger models without significant slowdown because its much faster if it can keep the entire model in memory. Small models range from 3-8 gigabytes, but a 70B parameter model will be 30-50 gigabytes.
- neop1x 2y agoI am running 70B models on M2 Max with 96 GB of RAM and it works very well. As HW evolves, it will become a standard