4 ms·
Now we're seeing decent 13-40B models (within ChatGPT 3.5 level). 4-5bit GPTQ quantizations are working pretty well, so these models are actually fitting on con
by mcbuilder 3y ago
Now we're seeing decent 13-40B models (within ChatGPT 3.5 level). 4-5bit GPTQ quantizations are working pretty well, so these models are actually fitting on consumer GPUs. So many new models on huggingface coming out everyday, it is hard to keep up with the foundation models. We are in a great time for regular users being able to play with LLMs. People are loading onto their consumer 4090 GPUs things models that would have been state of the art a couple years ago.
We're also plateauing with LLM performance, they aren't scaling past 200B it seems even with 1T+ token training (that's why they "copped out" and did MoE for ChatGPT)
I think LLM AI performance will likely stabilize, but I also think we'll easily get that 3 orders of magnitude in the next 5 years. So, maybe AI won't be super smart but it will be everywhere.
- famouswaffles 3y ago>(that's why they "copped out" and did MoE for ChatGPT) MoE outperform dense counterparts significantly after instruct tuning. https://arxiv.org/abs/2305.14705 https://arxiv.org/abs/2305.14705
- ianbutler 3y ago> We're also plateauing with LLM performance, they aren't scaling past 200B it seems even with 1T+ token training (that's why they "copped out" and did MoE for ChatGPT) Can you point out to anything other than speculation by Geohot here? I heard the same thing, but all of this has been circling the twitter sphere and I haven't seen any supporting research to back up this claim.
- mcbuilder 3y agoIn my opinion these are the two alternatives and we're left making educated guesses. I mean it's clearly rumors, but look at the scientific results on your bread and butter language modeling tasks you see any clear basic architecture wins. For instance look at hellaswag results, https://rowanzellers.com/hellaswag/ https://rowanzellers.com/hellaswag/. GPT-4 is impressive, but RoBERTa is not far behind and that's from 2019. It's from 2019 is the point I'm trying to drive home. T5, RoBERTA, Transformer XL, all old as hell (for ML/AI) architectures but still pretty top contenders. At this point I think we'd see more big and basic results at top conferences if we expect AI to keep scaling in "intellegence", but damn we're close to solving human language modeling in limited contexts. That's still huge, along with the advances in computer vision in the last 10 years and generative art, the rate of breakthroughs is incredible, but we're also going to be hitting brick walls now and again.
- famouswaffles 3y agoRoberta is tuned on Hellaswag so the comparison means nothing. There's a big difference in the uality of responses between 3.5 and 4, nevermind anything before that.