3 ms·
Where does the performance difference come from? And in what kind of processor & gpu? I didn't even know llama.cpp had a 32 bit option. For now I'm pretty suspi
by version_five 3y ago
Where does the performance difference come from? And in what kind of processor & gpu? I didn't even know llama.cpp had a 32 bit option. For now I'm pretty suspicious it's a fair comparison.
- tjake 3y agoThe default for `convert.py` is F32. This is just SIMD CPU comparison. Jlama uses the vector api in java20 but also better thread scheduling with work stealing and zero allocation.
- belfthrow 3y agoCould you link to some of the examples in your repo where you enforce the zero allocation? I don't see much reuse of the buffers, eg float buffers and there is quite a lot of array based heap allocation. Just for my own interest. Many thanks. Cool to see the use of the new vector api also.
- version_five 3y agoVery interesting, I'll watch for the quantized version.