3 ms·
A Java port of llama2.c that performs very close to C on large models. Llama 2 7B runs at a whooping 1.6 tokens/s.
by mukel 3y ago
A Java port of llama2.c that performs very close to C on large models. Llama 2 7B runs at a whooping 1.6 tokens/s.
- mike_hearn 3y agoHey man, awesome stuff. Surely any JIT compiler will struggle to vectorize something using IntStream.range, though? Looking at matmul, I'd not expect that to be auto-vectorized. The Panama API can be used to do a matmul vectorization, too bad it seems to never launch.
- mwcampbell 3y agoPanama is now in its third preview in the soon-to-be-released JDK 21: https://openjdk.org/jeps/442 https://openjdk.org/jeps/442 Is there any indication that it won't go from there to a final release soon?
- mike_hearn 3y agoThat's only for the FFI I think. The vector API has been incubated six times now and is waiting for Valhalla :(