4 ms·
Another big difference is quality of the results. Haven't tried myself but seen many complaints that it's nowhere near GPT-3 (at least for the 7B version). Corr
by tracyhenry 4y ago
Another big difference is quality of the results. Haven't tried myself but seen many complaints that it's nowhere near GPT-3 (at least for the 7B version). Correct me if I'm wrong!
- bestcoder69 4y ago13B feels on-par with the base non-instruction davinci. People might not realize how it was a bit trickier to prompt gpt3 when it first released.
- simonw 4y agoThat doesn't bother me so much. GPT-3 had instruction tuning, which makes it MUCH easier to use. Now that I've seen that LLaMA can work I'm confident someone will release an openly licensed instruction-tuned model that works on the same hardware at some point soon. I also expect that there are prompt engineering tricks which can be used to get really great results out of LLaMA. I'm hoping someone will come up with a good prompt to get it to summarization, for example.
- sp332 4y agoChatGPT had an estimated 20,000 hours of human feedback. That’s not going to be easy to replicate in an open source way.
- simonw 4y agoThat's the next level up from instruction tuning though: that was the RLHF stuff, which was essential to make ChatGPT useful and safe enough to expose to a wide audience. For a model running on my own laptop I'm OK taking more risks. I'd like it to be able to obey simple instructions like "Summarize this text" or "Extract the names of everyone mentioned in this article" - I don't care as much about the stuff ChatGPT has to get right.
- renewiltord 4y agoPerhaps we will provide feedback to open source Llama using ChatGPT. The cost to adjust the model is presumably what's hard?
- Taek 4y agoOpenAssistant has already collected 100,000 human feedback examples, estimated 5,000+ hours of human work via crowd sourced volunteers. Enough programmers want this badly enough that its going to happen. Inference at 8 GB and fine tuning at 24 GB, just like stable diffusion, on a 13B model.
- boredhedgehog 4y agoFrom what I gather, tuning a language model is like jury duty: The more eager someone is to volunteer, the less useful his input is going to be.
- Taek 4y agoI don't think the quality of the OpenAssistant data represents that idea at all.
- sebzim4500 4y agoThis sounds incredibly hard to believe, unless your concern is that advertisers will poison the well.
- vintermann 4y agoIs it that hard to believe? If you look at Stable Diffusion, I'd say a lot, maybe the majority of the volunteer effort was focused on anime girls and "realistic" pictures of anime girls (which amounts to young faces on adult bodies).
- koheripbal 4y agoThey had to chop the 30b and 65b models to 4-bit quantization which makes it significantly dumber.