6 ms·
We include a disclaimer later that researchers are debating whether it's possible to predict emergent capabilities. Wei has responded to that paper and others a
by elifland 3y ago
We include a disclaimer later that researchers are debating whether it's possible to predict emergent capabilities. Wei has responded to that paper and others at https://www.jasonwei.net/blog/common-arguments-regarding-emergent-abilities https://www.jasonwei.net/blog/common-arguments-regarding-eme... at I don't think it's clear who is right
- kromem 3y agoI think it's fairly clear Wei is right, given that the paper cited earlier really doesn't make the case that emergence isn't happening, it only makes the case that other measures exist by which improvement is linear, and thus not ALL metrics have emergent growth. As an aside, Wei's point at the end of that post about what happens with CoT's effectiveness at different model sizes is particularly brilliant.
- godelski 3y agoI actually disagree (and I'll also note that I don't like the term "emergent"). There's a few factors that are coupled with the analysis that matter here. Re: Metrics I don't think Wei is wrong about what he's said here but has responded to a rather weak form of the argument. It is correct that we, in the end, care about a binary distinction of getting the answer right vs wrong. But the issue here is that with hard metrics we have a very flat loss landscape and so there is little information being fed back to the network. You are perfectly capable of combining hard and soft metrics or even having the soft metrics decay or turn off after sufficient learning. I'm not aware of anyone that's explored this, but it is a natural hypothesis that we should have fairly high confidence that this would result in good results considering we already see smooth performance in soft metrics. Similarly it should be unsurprising that a hard metric has jumps. The larger models have a clear advantage here not just via data but because the number of parameters allows the model to fold/unfold the data through means that the smaller model couldn't and so the bigger question is about if the smaller model could learn such foldings given sufficient time. In other words, larger models can simply search a large solution space faster, so if it's takes a random hit to find a minima, the expectation that a large model finds one is going to be exceptionally higher than that for a small model. The idea here is also more abstract than his critique that cross-entropy on IPA transliterate still has a large kink because the ultimate question here is about how flat the loss landscape is and our expectation of stumbling upon non-flat regions. I simply would not expect smooth gains if our loss space (via metric or even via the problem itself) is flat with sparse optima. Re: UShape This is certainly a surprising phenomena and worthy of investigation. But I think it is also not clearly dismissed via the above framework. The losses need not be perfectly flat and as any good mathematician knows, a metric can lead you in the wrong direction if used wrong enough. I'm not saying this validates my claim above but rather that it doesn't invalidate it. If in fact the landscape is a very soft slope pointing away from a deep optima (think approaching a volcano, but a very soft grade) can result in this phenomena. This would be a very tough optimization problem but we do have many more opportunities to find the magma chamber with a large model. This also ties into chain of thought with essentially the same reasoning he gives. But I've always thought of chain of thought prompting as a bit of cheating. It's incredibly easy to introduce information leakage into a model and COT is often giving hints to your models. I do find a certain irony here given that he critiqued soft metrics earlier. I don't know if I'm wrong or right. But I do certainly think it is too early to dismiss the ideas. I do think we really do need to get more into (read advance) model interpretability to even approach these questions in a good way. I also think our community needs to stop shying away from math. > Focusing on metrics that best measure the behavior we care about is important because benchmarks are essentially an “optimization function” for researchers. I also want to address this, despite it being in the first Re. I will continue to rage against this idea, even if softly put in quotes. No metric is anywhere near close to the behaviors we actually care about. There are no metrics for quality of speech, visual fidelity, vocal realism, and so on. It's impressive that we've done so well when you dig into the metrics we use, but they were selected with care. At the same time, benchmarks are highly limited and especially in discussions of large models (language or vision) these metrics and benchmarks are showing their limitations. Simply due to the alignment of said metrics with desired behavior. We don't desire that the distribution that the LLM learned is indistinguishable from the distribution we used to train it (KL -> 0), but rather we desire that a LLM is able to write language well and perform complex tasks. These do relate, but they are not the same thing. It also makes it disingenuous to compare models with different training sets (such as comparing a JFT pretrained model to something else) due to the nature of what we're actually measuring (e.g. JFT may very well, and likely is, a better approximation of the desired object we're modeling with probability distributions and tuning it to have similar distributional properties to a subset distribution is far easier than training something to learn a distribution in the first place). The desired outcome is, as best we can tell, ineffable. It takes far more than a metric and a benchmark (or several) to quantify the performance of even simple models, let along these beautiful lovecraftian constructs.
- kromem 3y agoHow do you see CoT as "giving hints"? Isn't the whole point that the hints are being sourced from the model itself? Which is precisely why a more robust model has compounding gains on generating intermediate steps over a less robust model? And what are your thoughts on the various papers over the past year looking at transmitting capabilities from larger models to smaller models using synthetic data from the larger models? Rather than considering the 'emergent' (I agree not the best term) gains as a result of more parameters during operation, wouldn't this indicate that markedly better performance in larger models is the result of better network optimizations in those models developed during training, and that these optimizations can be successfully transferred to smaller parameter models by generating more optimized training data vs a broad basis training set?
- godelski 3y ago> How do you see CoT as "giving hints"? CoT prompting has you give examples. The exact sample from the paper is Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now? A (not CoT): The answer is 11 A (replace above with this for CoT): Roger started with 5 balls. 2 cans of 3 tennis balls each is 6 tennis balls. 5 + 6 = 11. The answer is 11. Both actually have demonstrations, which are hinting. The second A has stronger hinting because it is prompting the model to follow a certain form. There's a good example from this conversation a few weeks back[0]. I link to the parent and see their chat log vs mine. May also want to look at the context as other comments have relevant information. But you can see in mine that I'm being incredibly careful to not tell GPT anything other than it is wrong, then work towards the parent's method. The hinting is very subtle in these examples but they are enough to spoil the test. But to be clear, it depends on what we're testing. If we're testing if a model can get to a solution, then hint all you want as long as you don't explicitly tell it (CoT is often pretty close to this line though, as in the above example). But if you're testing how robust a model is, how "intelligent" it is by human standards, then CoT is cheating. The context matters. Essentially the more robust a model is the less prompt engineering it requires. In this sense humans are rather robust despite the large amount of disagreements we have. Of course you can also make the argument that humans' robustness is due to hinting and cultural priors which is why there's(?) a larger rate of miscommunication across cultures but that's a whole other can of worms discussion. Obviously the fact that you have to hint doesn't make it a bad tool, but it does tell you that you should be exceptionally cautious of results when you don't hint (maybe you don't know) or are outside its main wheelhouse. Which what this is is rather unknown, especially considering the sequential nature of the design. > And what are your thoughts on the various papers over the past year looking at transmitting capabilities from larger models to smaller models using synthetic data from the larger models? I'll address this and the next part here since they're tied together. Distillation is awesome. But it is also directly tied to these loss landscapes that I'm discussing. Your teacher model is essentially telling you "hey, this way" because it already has explored the landscape. A teacher model doesn't even have to be bigger than the student model, just better. Now with generative models it's important to remember that they are also classifiers (and your classifier is secretly a generative model[1]. The arguments shouldn't be surprising if you have a deep understanding, though might result in you feeling 1) dumb because you didn't realize it a priori or 2) overly confident because "of course" and you forgot your a priori predisposition. The tyranny of good ideas). In the current state of the art typically GANs still reign in regards to actual fidelity (tricking humans) and sampling speed but diffusion's big win is diversity. This is due to diffusion models approximating the density function of the target distribution (what we're trying to learn)[2]. One way to think about this is the density of the distribution as a whole. Imagine a solid red circle but some parts are more red than others and some chunks are just not red at all! The better model should be more uniformly red (our target is all red) and so sampling from this we are going to more evenly sample from areas that are underrepresented. Pretty awesome! Of course we have to be careful and there is a lot of nuance too. You're not going to teach the student model something the teacher hasn't learned, although you might see new behavior since the student can better be primed to learn something the teacher didn't (yeah, this gets messy super fucking fast). There's also the danger of tightening the distribution and locking you out of learning things you want to learn. But this should help explain why even given the same training data we're starting to see bigger improvements, because we often don't actually care about fidelity (especially since transformers fucking love augmentations and noise injection is rather an important augmentation). We really care more about distributions. Of course, if you wanted to be really tricky (but this would likely be very expensive) you could monitor the density of the student model and have the teacher intelligently increase sampling density to these regions. Really the tldr here is thinking about the geometry of our models (we are assuming manifolds so we can leverage this, even if the assumption isn't absolutely correct). I tell students and by lab members that you don't need math to train good models, but you do need math to know why your models are wrong. It's also incredibly important to discuss our limitations because that's like knowing where our low density regions are and where we need to over invest in sampling to create better models. Essentially the nuance is incredibly important and we shouldn't shy away from it. Okay sorry, this got a bit ranty. Being terse is hard. [0] https://news.ycombinator.com/item?id=36307880 https://news.ycombinator.com/item?id=36307880 [1] https://openreview.net/forum?id=Hkxzx0NtDB https://openreview.net/forum?id=Hkxzx0NtDB [2] Note that the real world breaks a lot of our assumptions in ML. Things aren't always distributions, let alone i.i.d. Nor does data always lie on a a manifold and realistically it most certainly does not. Also note that the diffusion model is a reduction or what we'd call an approximate density function. A VAE is a clearer example because we typically do dimensional reduction but this is not necessary. Diffusion is because we don't have bijective maps. For exact density you're going to look at Normalizing Flows or Autoregressive models (don't confuse) or NODEs (I kinda call these all the same thing though NFs though tbh since they're isomorphic transformations). GANs are called implicit density estimators since you don't actually learn the density function but rather just a distribution that has similar sampling features.