3 ms·
In my opinion these are the two alternatives and we're left making educated guesses. I mean it's clearly rumors, but look at the scientific results on your brea
by mcbuilder 3y ago
In my opinion these are the two alternatives and we're left making educated guesses. I mean it's clearly rumors, but look at the scientific results on your bread and butter language modeling tasks you see any clear basic architecture wins. For instance look at hellaswag results, https://rowanzellers.com/hellaswag/ https://rowanzellers.com/hellaswag/. GPT-4 is impressive, but RoBERTa is not far behind and that's from 2019. It's from 2019 is the point I'm trying to drive home. T5, RoBERTA, Transformer XL, all old as hell (for ML/AI) architectures but still pretty top contenders.
At this point I think we'd see more big and basic results at top conferences if we expect AI to keep scaling in "intellegence", but damn we're close to solving human language modeling in limited contexts.
That's still huge, along with the advances in computer vision in the last 10 years and generative art, the rate of breakthroughs is incredible, but we're also going to be hitting brick walls now and again.
- famouswaffles 3y agoRoberta is tuned on Hellaswag so the comparison means nothing. There's a big difference in the uality of responses between 3.5 and 4, nevermind anything before that.