3 ms·
I don't know about these large models but I saw on a random HN comment earlier in a different topic where someone showed a GPT-J model on CPU only: https://gith
by adeon 4y ago
I don't know about these large models but I saw on a random HN comment earlier in a different topic where someone showed a GPT-J model on CPU only: https://github.com/ggerganov/ggml https://github.com/ggerganov/ggml
I tested it on my Linux and Macbook M1 Air and it generates tokens at a reasonable speed using CPU only. I noticed it doesn't quite use all my available CPU cores so it may be leaving some performance on the table, not sure though.
The GPT-J 6B is nowhere near as large as the OPT-175B in the post. But I got the sense that CPU-only inference may not be totally hopeless even for large models if only we got some high quality software to do it.
- generalizations 4y agoThere's also the Fabrice Bellard inference code: https://textsynth.com/technology.html https://textsynth.com/technology.html. He claims up to 41 tokens per second on the GPT-Neox 20B model.