4 ms·
Google's Pathways Language Model and Chain-of-Thought
- PaulHoule 4y agoI've talked about structural deficiencies in earlier language models, this one seems to be doing something about them.
- vackosar 4y agoSounds interesting! Would you link to that or describe them here? Thanks!
- PaulHoule 4y agoA very simple one is "can you write a program that might never terminate?" If a neural network does a fixed amount of computation and that is that it is never going to be able to do things that require a program that may not terminate. There are numerous results of theoretical computer science that apply just as well to neural networks and other algorithms even though people seem to forget it. Another is "can an error discovered in late stage processing be fed back to an early stage and be repaired?" That's important if you are parsing a sentence like Squad helps dog bite victim. It was funny because I saw Geoff Hinton give a talk in 2005, before he got super-famous, and he was talking about the idea that led to deep networks and he had a criticism of "blackboard" systems and other architectures that produced layered representations (say the radar of an anti-aircraft system that is going to start with raw signals, turn those into a set of 'blips', coalesce the 'blips' into tracks, interpret the tracks as aircraft, etc.) Hinton said that you should build the whole system in an integrated manner and train the whole thing working end-to-end and I thought "what a neat idea" but also "there is no way this would work for the systems I'm building because it doesn't have an answer for correcting itself.
- cygaril 4y agoYou're assuming here that there are discrete stages that do different things. I think a better way to conceptualise these deepnets is that they're doing exactly what you want - each layer is "correcting" the mistakes of the previous layer.
- PaulHoule 4y agoMost "deep" networks are organized into layers and information flows in a particular direction although it doesn't have to be that way. Hinton wasn't saying we shouldn't have layers but that we should train the layers together rather than as black boxes that work in isolation. Also, when people talk about solving problems they talk about layers, layers play a big role in the conceptual models people have for how they do tasks even if they don't really do them that way. For instance in that ambiguous sentence somebody might say it hinges on whether or not you think "bite" is a verb or a noun. (Every concept in linguistics is suspect, if only because linguistics has proven to have little value for developing systems that understand language. For instance I'd say a "word" doesn't exist because there are subword objects that depend like a word "non-" and phrases that behave like a word (e.g. "dog bite" fills the same slot as "bite")) Another ambiguous example is this notorious picture https://www.livescience.com/63645-optical-illusion-young-old-woman.html https://www.livescience.com/63645-optical-illusion-young-old... which most people experience as "flapping" between two states. Since you only see one at a time there is some kind of inhibition between the two states. Who knows how people really see things, but if I'm going to talk about features I'm going to say that one part is the nose of one of the ladies or the chin of the other lady. Deep networks as we know it have nothing like that.
- space_fountain 4y agoI'm by no means an expert, but a lot of choices machine learning algorithms make are more about training parallelization than anything. In many ways it feels like something like a recursive neural network or some architecture even more weird should be better for language, but in practice it's harder to train an architecture that demands each new output depend on the one before. Introducing dependencies on prier output typically kills parallelization. Obviously this is less of a problem for say a brain that has years of training time, but more of problem if you want to train one up in much less time using compute that can't do sequential things very quickly
- phoe18 4y agoThe article quotes the cost as roughly 10B$ in the first paragraph. Likely a typo? They quote 10M$ in a later paragraph.
- simulate-me 4y agoThe amount of capital needed to train these high-quality models is eye watering (not to mention the costs needed to acquire the data). Does anyone know of any well capitalized startups exploring this space?
- vackosar 4y agoCorrection! Cost is around $10M not $10B.
- lumost 4y agoOpenAI would be the best example. However these large language models also have limited business value today, making an startup a speculative bet that the team will beat Google/FB/AI/Academics at making a language model and find a viable business model for the resulting model. I'd take one of those bets or the other, both are tough to pull off. Considering that the first task of such a startup would be to hand ~100-500MM to a hardware or cloud vendor I'd be hesitant to invest as an investor.
- simulate-me 4y agoI agree 100%, but I think viable businesses will begin to emerge especially as these large models move from text to images (and eventually to video and 3d models). If the examples shown of DALL-E 2 are indicative of its quality, then a large number of creative jobs could be replaced with a single "creative director" using the model. But the high entry cost just to attempt to train such a model will likely remain a hurdle until more business value is proven.
- lumost 4y agoaye - I suspect the other concern is hat the high entry costs can quickly lead to a "second mover" advantage. The first team spends all the money doing the hard R&D and the second team implements a slightly better version for a fraction of the money.
- 4y ago
- vackosar 4y agoCorrection!! The model costed around 10M not 10B! Thanks for raising that. Mistake during copying from the second slide :(
- imranq 4y ago$10M for a bag of numbers (i.e the learned weights of the model matrices)