3 ms·
Tell us how it goes! Try different numbers of layers if needed. A good place to dig for prompt structures may be the 'text-generation-webui' commit log. For ex
by rain1 3y ago
Tell us how it goes! Try different numbers of layers if needed.
A good place to dig for prompt structures may be the 'text-generation-webui' commit log. For example https://github.com/oobabooga/text-generation-webui/commit/334486f527bc97f61eb3264def4e03a0dab9b369 https://github.com/oobabooga/text-generation-webui/commit/33...
- int_19h 3y agoI tried llama-65b on a system with RTX 4090 + 64Gb of DDR5 system RAM. I can push up to 45 layers (out of 80) to the GPU, and the overall performance is ~800ms / token, which is "good enough" for real-time chat.