5 ms·
It seems still unclear how much quality loss there is compared to the best models. What's really needed is systematic evaluation of the output quality, but that
by blueblimp 4y ago
It seems still unclear how much quality loss there is compared to the best models. What's really needed is systematic evaluation of the output quality, but that's tricky and relatively expensive (compared to automated benchmarks), so I understand why it hasn't happened yet.
Edit: I just tried it with a single task of my own (that I've successfully used with ChatGPT and Bing) and it flubbed it horribly, so this model at least is noticeably inferior to the SOTA, which is not surprising given how small it is.
- yunyu 4y agoI assume you haven't tried Alpaca (which hasn't been released), only Llama. See the instruction fine tuning section in the article.
- karmasimida 4y agoThey currently only supports a single input/response format of input right? Multi turns will be more challenging to handle. I am optimistic for 30B or 66B to catch up with OpenAI, but 7B is unlikely to have the same quality.
- enlyth 4y agoHere is a sample chatbot style conversation I had locally with LLAMA 30B 4bit quantization on a 3090 RTX (generation speed 5 tokens per second): https://i.imgur.com/V4lzLz7.png https://i.imgur.com/V4lzLz7.png It does quite well at simple reasoning. More complex stuff, it does struggle sometimes, but I am impressed by the output and the fact that it even runs on my computer. I used the oobabooga repository on Ubuntu 22.
- plaba417 4y agoI have had a few questions where it gave a completely nonsensical answer, but by fiddling with the parameters, I've gotten more usable output. Have you tried lowering the noise and increasing top_p?