5 ms·
This, again, sparks the "is this general ai?" question, which often results in low quality, borderline-flaming content... My take: the point of this paper isn'
by inductive_magic 4y ago
This, again, sparks the "is this general ai?" question, which often results in low quality, borderline-flaming content... My take:
the point of this paper isn't "here, we solved general intelligence". It's "look, multi modal token prediction is a sound iteration". Look at the scale of the model in comparison to, say, gpt-3: this is a PoC, they didn't bother scaling it, because we've already seen where scaling these mechanisms leads.
What I would love to know is what kind of architectures deepmind et al are playing with in-house. Token prediction is a promising avenue, but it's more of a language that an intelligent agent may operate in, opposed to the self-sufficient structure of the intelligent agent itself -- the symbolic system that implements algos like gato. If that symbolic system will be the result of a generator-function, that generator function won't be token prediction by trade. I mean, maybe somewhere in the deep depths of a multi modal model, intelligent structure may emerge, but that would be a very weird byproduct.
- Barrin92 4y ago>but it's more of a language that an intelligent agent may operate in, opposed to the self-sufficient structure yes, this kind of functional intelligence seems distinct from an actual living entity, which is the thing that uses subordinate functions to pursue goals and has some interior state, motivations and some sort of architecture. To reduce intelligence to tokens predicting more tokens is kind of like saying f(x), just solve for intelligence. When prediction itself is only partially what intelligent systems are about. Agent is a very important word because it's accurate ("a means or instrument by which a guiding intelligence achieves a result") And it's the latter I think we ought to be after when talking about 'general ai'.
- jawarner 4y agoIt’s possible that in serving the function of prediction, the model forms a complex internal representation akin even to goals, motivations, etc. It is true that DL architectures are not explicitly designed to do this, not yet anyway. But my point is that the task of prediction can give rise to such architectural patterns. According to Karl Friston’s Free Energy Principle, biological brains serve the purpose of predicting the value of different actions available.
- hooande 4y agothis assumes there is a finite list of available actions. for example, a primate has to see that "sharpen a stick to use as a spear" is an option, and add it to the list
- Barrin92 4y agoI agree with that, I think it's even necessarily true in natural intelligence which after all emerged by some means spontaneously. But scientifically I think it is a big problem because it does not supply us with a theory or systematized knowledge of the mind or intelligence. I think you could even imagine say, why not just make a primordial soup simulation, insert some DNA, and crank the speed up, intelligence is just a byproduct somewhere in there in a physics simulation. don't even bother with such details as neural nets. Scientifically this is unsatisfying but also if for some reason this turns out to be an engineering dead-end we have a big hole where a concrete theory of intelligence should be, with its components, mechanisms and so forth. And sadly I think this is still the weakest link in AI. To me it seems a little bit like if you trained architects instead of having a theoretical basis for architecture, you just showed them every building in existence and sent them to work. It may very well work, but if it didn't you have a problem. And even if it did, you'd still want to have an understanding of why it works.
- Jack000 4y agoSometimes there is just no reductive theory for emergent phenomena. For example, in the case of ant routing we have a good understanding of the behavior of individual ants, and we can observe the intelligent behavior of the colony as a whole. One could ask for a concrete theory of how the micro scale behavior of each ant leads to the macro scale behavior of the colony, but it doesn't exist. If we had enough working memory to hold all the ants in our minds, we'd see that both micro and macro scale behaviors are different aspects of the same system, there's no "missing link" in the chain of explanations, and the dichotomy is only due to our own cognitive limitations. There are a lot of problems that do not have analytical solutions, like the N-body problem. With these problems all you can do is numerical simulation, which is pretty much what modern machine learning is.
- sva_ 4y ago> because we've already seen where scaling these mechanisms leads. In the case of GPT-3, scaling seemed to continuously improve results, they just kinda ran out of data. Are you implying this must be the same for this model? Or were you intending to say something different that I didn't see?
- atty 4y agoIt explicitly says in the article that the reason they capped it at ~1B parameters is because that’s the current limit of what they can achieve and hit their latency requirements? It has nothing to do with not being interested in scaling it further, as far as I’m aware.
- akomtu 4y agoImo, the current "the bigger the better" trend is limiting further progress. Intelligence finds the simplest model that explains all data, e.g. a small set of equations or rules, while the today's wannabe-AI models are trying to remember all the data in a fuzzy lookup table.
- atty 4y agoI think there’s two parts to this. The first is that there’s a minimum number of parameters required to do an arbitrary task. For any sufficiently complex task (image recognition, large language models, etc), it’s not clear how to find that lower bound. And the bound probably depends on the model chosen. But we do know that the more parameters you add, the more complex of a function you can learn. On the other hand, for many reasons (energy, training/inference deployment complexity, latency, even some sense of model elegance I suppose), we don’t want to massively increase the number of parameters unnecessarily. But again, I don’t think we have great methods to estimate the “ideal” minimum number of parameters for a model to achieve its goals. And what we keep finding is that if you increase the number of parameters, and you increase the training corpus, the model gets more accurate, more impressive, and it’s not stopping. So while I definitely agree that size for size’s sake is wasteful, I also don’t think we necessarily even know how to define “wasteful” for things like large language models right now.
- jona-f 4y agoWell, while this is very interesting research, the title of the paper is a bit misleading. They way I understand the paper, their network is specialized in multiple domains, but that doesn't make it a generalist. Clearly there is a lot of potential, but skimming through quickly, I didn't find much about out of domain data. Generally, I think nomenclature in the ai world is horrible. A mishmash of different academic disciplines and hyped keywords. Attention isn't really attention (more like association), generative adversarial networks aren't really adversarial (they are designed to work together in the end), ...