3 ms·
Well, I was looking for any excuse to upgrade to AM5, so I think this qualifies. Question: if I already have a previous version .llamafile, is there a way to r
by fbdab103 3y ago
Well, I was looking for any excuse to upgrade to AM5, so I think this qualifies.
Question: if I already have a previous version .llamafile, is there a way to repack the model blob with the new llamafile executable? That is, say I have mixtral_llamafilev6 promote to mixtral_llamafilvev7?
From the release page
- Prompt evaluation now goes much faster on CPU. For example, f16 weights on Raspberry Pi 5 are now 8x faster. These new optimizations mostly apply to F16, BF16, Q8_0, Q4_0, Q4_0, and F32 weights. Depending on the hardware and weights being used, we've observed llamafile-0.7 going anywhere between 30% to 500% faster than llama.cpp upstream.
...
- Support for AVX512 has been introduced. Owners of CPUs like Zen4 can expect to see 10x faster prompt eval times.
- formerly_proven 3y agoYes, they're just ZIP files that also happen to be actually portable executables, ehm, excuse me, αcτµαlly pδrταblε εxεcµταblες. https://github.com/Mozilla-Ocho/llamafile?tab=readme-ov-file#creating-llamafiles https://github.com/Mozilla-Ocho/llamafile?tab=readme-ov-file...
- fbdab103 3y agoSure enough. Pop open the archive and the gguf is right there. There is even a `.args` file where you could instruct it to launch a different model name.