4 ms·
The author makes very compelling arguments within the context of known tools/techniques that we have today. However, they conveniently sidestep the possibility
by mrshadowgoose 3y ago
The author makes very compelling arguments within the context of known tools/techniques that we have today. However, they conveniently sidestep the possibility of the development of new tools and techniques. This sidestep is overtly stated under the "Active Learning Is Hard & Fundamental" section. And that entire section basically boils down to "I personally believe this is hard and will take a long time."
The only reason we have things like ChatGPT today is because of the surprise development of the transformer. Nobody understands how intelligence fundamentally works, so we literally have no ability to predict when the next groundbreaking architectural development will occur.
And self-improvement concerns are mostly orthogonal to the main risk of AI: the development of AGI. Even a mediocre AGI that is only "decent" at tasks, can be weaponized as a collective superintelligence.
- turkeygizzard 3y agoThank you for articulating this. I remember similar problems and arguments arising after RNNs and CNNs became massively successful. People argued that training larger models would be infeasible for several reasons that all were made moot by Attention Is All You Need. Somebody seems to always figure out a new approach
- hackernewds 3y agoAttached is the Attention is all you need publication. https://arxiv.org/abs/1706.03762 https://arxiv.org/abs/1706.03762 Why was this revolutionary though?
- uh_uh 3y agoPrevious approaches like LSTM struggled learning long-term dependencies. The transformer improved on this greatly.
- yinser 3y agoTo add another anecdote to your question: the transformer became a part of the first context aware embedding model GPT-1. Not to say it couldn’t be done with another tool but it was first done with a transformer. Previous embedding models like word2vec, GloVe and fasttext were not contextually embedding and didn’t give you a language graph that would then go on to support a language model capable of “understanding” what you were saying or asking for.
- jimsimmons 3y agoGP is wrong. Attention is all you need paper just proposed an AR model that didn’t have to be trained step by step. The scaling happened later in BERT and GPT and OpenAI’s scaling work
- abetusk 3y agoI'm not sure I understand it well enough to say but watching a video on it [0] I think there were a few key points: * "Attention is all you need" introduced positional encoding which allows you to keep context of the word, allowing for more complex translation (and thus generative/chatgpt like tasks?) because words now have context relative to each other. Contrast this with "bag of words" models that only tells you whether the word is present or not. * I don't quite understand why but transformers (which "AiaYN" introduced) can be made fully parallel, compared with the RNN/LSTM networks which has to be serial per token. Fully parallel allows for GPU optimization, which means you can take advantage of Moore's law for training. I'm always a bit suspicious when people claim a breakthrough of this sort. There's no doubt that better algorithms give better results but how much is due to just faster computers, cheaper compute, memory, etc. [0] https://youtu.be/S27pHKBEp30 https://youtu.be/S27pHKBEp30
- Certhas 3y agoFirst of all, the article argues that you need a major breakthrough, arguably attention was such a breakthrough? That said, this doesn't really seem all that comparable. The article points out very fundamental properties of all the diverse current approaches: They are tightly data constrained. You either need to cheap simulation or massive real world data. That's not an arcane technical point.
- scythmic_waves 3y agoYour comment confuses me. You say the author > conveniently sidestep[s] the possibility of the development of new tools and techniques. I don't believe they did this at all. Here is their summary of their third point: > To automatically construct a good dataset, we require an actionable understanding of which datapoints are important for learning. This turns out to be incredibly difficult. The field has, thus far, completely failed to make progress on this problem, despite expending significant effort. Cracking it would be a field-changing breakthrough, comparable to transitioning from alchemy to chemistry. That does not read to me as conveniently sidestepping the possibility of new tools. Rather, it is acknowledges that to overcome the problem we NEED a new tool, aka a breakthrough. I also disagree with the framing of the third section as: > And that entire section basically boils down to "I personally believe this is hard and will take a long time." They provide both theoretical and empirical evidence of their claim. I find it was well argued, and I'm inclined to agree. Every scenario I can think of with a runaway superintelligence requires a way to automatically improve datasets as part of the learning process.
- mrshadowgoose 3y agoFrom one of the final paragraphs: "Also, they aren’t getting more-solved over time: we’ve made little-to-no progress on any problem of this sort in the last decade, certainly not the reliable improvements of the sort we’ve seen from supervised learning. This indicates that a breakthrough is needed — and that it is unlikely to be close." Past lack-of-progress is not an indicator of future lack-of-progress. "unlikely to be close" is unknowable, and is just the author's gut feeling.
- scythmic_waves 3y ago> Past lack-of-progress is not an indicator of future lack-of-progress. Past lack-of-progress is not proof of future lack-of-progress. But it's most definitely an indicator. > "unlikely to be close" is unknowable, and is just the author's gut feeling. Again I'm sorry but I have to disagree with you here. The very next paragraph reads: > Something that would change my mind on this is if I saw real progress on any problem that is as hard as understanding generalization, e.g. if we were able to train large networks without adversarial examples. Basically, the author has identified a class of problem. They are claiming that little to no progress has been made on any problem in that class. So it's not just that no progress has been made on this specific problem, but that we are stuck on all problems of this type. Thinking of it in that way, I do not find it unreasonable to say we are "unlikely to be close" to a solution. If you have a CS background, I'll make an analogy to computability classes: It's like the author is saying this is an NP hard problem. We have made no progress on any NP hard problem. I think it's reasonable to say that we are unlikely to be close to solving a particular NP hard problem because we've made no progress on any NP hard problem. You're free to disagree with the author's premises. E.g. "this problem is not like those other problems" or "we actually have made progress on those other problems". But I think that the conclusions are reasonable given the premises.
- cyanydeez 3y agoAdvanced technology is indistinguishable from magic. Your critique is: has the author not considered magic?