3 ms·
(didn't read because paywall) From experimenting with Falcon-40b-instruct the base model, outputs are acceptable for very simple tasks though still dwarfed by
by courseofaction 3y ago
(didn't read because paywall)
From experimenting with Falcon-40b-instruct the base model, outputs are acceptable for very simple tasks though still dwarfed by closed-source LLMs in usefulness.
The thing stopping me from experimenting further is how expensive inferences are - 40b to takes 8-10mins per 200 token generation, loaded in 4-bit with transformers on a colab A100. The model weighs in around 27GB on VRAM with these settings.
Has anyone gotten faster inferences with the 40b model?
- sliken 3y agoAnyone tried Falcon 40b on a Apple M2 max or M2 ultra? How many tokens per second?
- flangola7 3y agoWhy is the inference so sluggish?
- jxy 3y ago1 to 2 tokens per second on CPU, depending on the CPU/RAM, thanks to GGML. See discussions: https://github.com/ggerganov/ggml/pull/231 https://github.com/ggerganov/ggml/pull/231 And the code: https://github.com/jploski/ggml/tree/falcon40b https://github.com/jploski/ggml/tree/falcon40b