3 ms·
I don’t think we’re close to super human intelligence in the colloquial sense. ChatGPT scrapes all the information given, then predicts the next token. It has
by spacephysics 3y ago
I don’t think we’re close to super human intelligence in the colloquial sense.
ChatGPT scrapes all the information given, then predicts the next token. It has no ability to understand what is truthful or correct. It’s as good as the data being fed to it.
To me, this is a step closer to AGI but we’re still far off. There’s a difference between “what’s statistically likely to be the next word” vs “despite this being the most likely next word, it’s actually wrong and here’s why”
If we say, “well, we’ll tell chatgpt what the correct sources of information are” that’s no better really. It’s not reasoning, it’s just a neutered data set.
I imagine they need to add something like chatgpt 4 has with live internet models or something else to get the next meaningful bump
I don’t recall who said it, but a similar thread had a researcher in the field express that we have squeezed far more juice than expected from these transformer models. Not that new progress in this direction can be made, but it seems like we’re approaching diminishing returns
I believe the next step that’s close is to have these train on less and less horsepower. If we can have these models run on a phone locally, oh boy that’s gonna be something
- famouswaffles 3y agoGPT's already forgo the surface level statistically most likely next word for words that are more context appropriate. That's one of the biggest reasons they are so useful. The truth is that functionally/technically, there's plenty left to squeeze. The bigger issue is that we're hitting a wall economically.
- EGreg 3y agoHow do they do that? No one seems to have a real explanation of what OpenAI actually did to train it
- mindwok 3y agoInformation on how they trained it nonwithstanding, there’s clearly more than just statistically appropriate words going on because you can ask it to create completely new words based on rules you define and it will happily do it.
- feanaro 3y agoWell yes -- it's not words, it's tokens, which are smaller than words.
- famouswaffles 3y agoIt's pretty much just scale, either via Dataset size or parameter size. Before GPT-4, the general SOTA model was not in fact from Open AI (Flan-PaLM from Google). The attention from GPT-4 is a little different (probably some kind of flash attention) so that memory requirements for longer contexts are no longer quadratic. But there's nothing to suggest the intellectual gains from 4 isn't just bigger scale. Google could have made a 4 equivalent I'm sure. It's not like there wasn't a road to take. We already knew 3 was severely undertrained even from a computer optimal perspective. And then of course, you can just train on even more tokens to get them even better.
- EGreg 3y agoHow do you know it’s pretty much just scale? “Open”AI has been pretty tight-lipped about the details of its training and merely “claims” it was scale. It hired a ton of humans to train the model in little ways. If that’s the “scale” you’re talking about then it’s humans all the way down: https://www.forbes.com/sites/kenrickcai/2023/04/11/how-alexandr-wang-turned-an-army-of-clickworkers-into-a-73-billion-ai-unicorn/ https://www.forbes.com/sites/kenrickcai/2023/04/11/how-alexa...
- famouswaffles 3y agoIt's just been scale up to before GPT-4. Open AI hasn't said anything specific about 4 so feel free to think there's something major going on if you want but there's basically nothing to support that.
- EGreg 3y agoHow do we know it’s just been scale up to before GPT-4? As I said, OpenAI hasn’t told us why ChatGPT is so much better than other models (Bloom, LaMDA, LLaMa) and yet we know that they employ thousands of people to do RLHF, including for the coding models: https://www.semafor.com/article/01/27/2023/openai-has-hired-an-army-of-contractors-to-make-basic-coding-obsolete https://www.semafor.com/article/01/27/2023/openai-has-hired-... Doesn’t quite sound like it’s “just scale”. I asked ChatGPT about its training and corpus and it explicitly disavows having that information.
- fnordpiglet 3y agoattention is all you need (Well and crap tons of GPUs and training data)
- firecall 3y ago> ChatGPT scrapes all the information given, then predicts the next token. It has no ability to understand what is truthful or correct. It’s as good as the data being fed to it. That is precisely true of Humans as well though! :-)
- ChatGTP 3y agoThe parrot is strong.