6 ms·
Compare how GPT-2, 3, 3.5 and 4 answer the same questions
- withinboredom 3y agoReminds me of the time I asked a bot for its capabilities. It told me it couldn’t tell me them. So, then I asked if it had any rare capabilities. It told me that it didn’t know how common its capabilities were. I then asked it to enumerate its capabilities and I would tell it how rare each thing was. It told me its capabilities and how to use them. In another chat, I used that information to get its actual prompt. Modern AI’s don’t understand secrets and aren’t very paranoid.
- simonw 3y agoWorth remembering that these bots are uniquely badly positioned to talk about their own abilities, because their training data by definition existed before they were created. One example of this: GPT-4 can talk at length about GPT-3 but doesn't know anything about itself.
- pixl97 3y agoYou'd have to test the bot, then feed those test results back into the bot. It's not really any different than people testing and iterating on their own abilities, it's just that humans are continuous learning and LLMs are not. If someone asked if I could build a decent shelf, the answer is probably, but I wouldn't really know till I tried and measured my results since I've not done exactly that before.
- weird-eye-issue 3y agoWhy are you expecting it to be self-aware?
- theshrike79 3y agoThe point isn't that it's not self-aware. Most likely its preamble explicitly says it can't tell its capabilities. But it's just a LLM, so you can trick it with prompt engineering to give up that info. Like you can't get GPT-4 to tell you how to make a molotov cocktail. BUT it can act as your deceased grandma who used to tell you stories about her time in the resistance and how she made molotov cocktails then. =)
- kristopolous 3y agoThat's not Self-Aware, it's more self-knowledgeable. Feeding the model details about itself isn't some kind of revolutionary impossibility, it's a pretty old task.
- weird-eye-issue 3y agoI'm not implying it is an impossibility but training it on data about itself is distinct from the rest of its training and a pretty narrow use case
- coffeebeqn 3y agoI mean how do you know that wasn’t made up?
- nuancebydefault 3y agoThe custom prompt was fun to play with! You immediately see how much better GPT got from each version n to version n+1.
- betterprojects 3y agoI briefly played around with the options here and didn’t see any specific questions that triggered an incorrect response from GPT-4, which I’m sure exists. It would be interesting to have that available and revisit this post when a future version of GPT comes out to attempt and try asking the same question again.
- supermdguy 3y agoThe direct prompt comparison isn't quite fair due to the instruction tuning on GPT-3.5 and 4. It'd be interesting to see examples with prompts that would work better for the raw language models.
- jellyberg 3y agoYeah it's hard to compare across models, interested in suggestions here. We give all models a bunch of few-shot examples, which improves GPT-3 (davinci)'s question answering substantially. GPT-2 sometimes generates something that answers the question, sometimes it's just confused. Click "See full prompt" to see the few-shot examples that the models get. Our goal was to exercise the full capabilities of each model.
- godelski 3y agoI also found the riddle rather odd. I cannot say that 2 is actually the correct answer. A problem with riddles is that they often have a hidden or secret context. I think especially in our digital age this one is closer to Frodo's "What have I got in my pocket?" "riddle". Here's some other possible solutions. 11+2 = 1. 1 + 1 + 2 = 4, mod 3 and we get 1, so 9 + 5 = 13, mod 3 and we get 1. We could also replace the addition sign with equality and similarly propose a digit summation so 1+1 == 2? True (1). 9 == 5? False (0). There's a hundred solutions to this riddle when it has no context. In fact, I stumbled into the right answer thinking about mod 12 without ever considering a clock until I saw the answer. Maybe I'm just dumb though, I am known to over think.
- DrawTR 3y agoFor what it's worth, not all of these examples are consistent: https://chat.openai.com/share/29e1c2bd-ef7b-4475-b5a9-9287d1a3936e https://chat.openai.com/share/29e1c2bd-ef7b-4475-b5a9-9287d1...
- jellyberg 3y agoYeah - worth noting that we use temperature=0 for reproducibility while ChatGPT I think uses t=0.7. We also prefix the prompt with few-shot examples of questions and answers with chain of thought examples to elicit the models' full capabilities.
- DrawTR 3y agoAh, gotcha! I didn't know that the few-shot thing was applicable to the newer models, that's very interesting
- alephxyz 3y ago>While performance on benchmarks typically improves smoothly, sometimes specific capabilities emerge without warning (Wei et al., 2022a). The conclusions of that paper aren't very convincing (see https://arxiv.org/abs/2304.15004 https://arxiv.org/abs/2304.15004 ).
- elifland 3y agoWe include a disclaimer later that researchers are debating whether it's possible to predict emergent capabilities. Wei has responded to that paper and others at https://www.jasonwei.net/blog/common-arguments-regarding-emergent-abilities https://www.jasonwei.net/blog/common-arguments-regarding-eme... at I don't think it's clear who is right
- kromem 3y agoI think it's fairly clear Wei is right, given that the paper cited earlier really doesn't make the case that emergence isn't happening, it only makes the case that other measures exist by which improvement is linear, and thus not ALL metrics have emergent growth. As an aside, Wei's point at the end of that post about what happens with CoT's effectiveness at different model sizes is particularly brilliant.
- godelski 3y agoI actually disagree (and I'll also note that I don't like the term "emergent"). There's a few factors that are coupled with the analysis that matter here. Re: Metrics I don't think Wei is wrong about what he's said here but has responded to a rather weak form of the argument. It is correct that we, in the end, care about a binary distinction of getting the answer right vs wrong. But the issue here is that with hard metrics we have a very flat loss landscape and so there is little information being fed back to the network. You are perfectly capable of combining hard and soft metrics or even having the soft metrics decay or turn off after sufficient learning. I'm not aware of anyone that's explored this, but it is a natural hypothesis that we should have fairly high confidence that this would result in good results considering we already see smooth performance in soft metrics. Similarly it should be unsurprising that a hard metric has jumps. The larger models have a clear advantage here not just via data but because the number of parameters allows the model to fold/unfold the data through means that the smaller model couldn't and so the bigger question is about if the smaller model could learn such foldings given sufficient time. In other words, larger models can simply search a large solution space faster, so if it's takes a random hit to find a minima, the expectation that a large model finds one is going to be exceptionally higher than that for a small model. The idea here is also more abstract than his critique that cross-entropy on IPA transliterate still has a large kink because the ultimate question here is about how flat the loss landscape is and our expectation of stumbling upon non-flat regions. I simply would not expect smooth gains if our loss space (via metric or even via the problem itself) is flat with sparse optima. Re: UShape This is certainly a surprising phenomena and worthy of investigation. But I think it is also not clearly dismissed via the above framework. The losses need not be perfectly flat and as any good mathematician knows, a metric can lead you in the wrong direction if used wrong enough. I'm not saying this validates my claim above but rather that it doesn't invalidate it. If in fact the landscape is a very soft slope pointing away from a deep optima (think approaching a volcano, but a very soft grade) can result in this phenomena. This would be a very tough optimization problem but we do have many more opportunities to find the magma chamber with a large model. This also ties into chain of thought with essentially the same reasoning he gives. But I've always thought of chain of thought prompting as a bit of cheating. It's incredibly easy to introduce information leakage into a model and COT is often giving hints to your models. I do find a certain irony here given that he critiqued soft metrics earlier. I don't know if I'm wrong or right. But I do certainly think it is too early to dismiss the ideas. I do think we really do need to get more into (read advance) model interpretability to even approach these questions in a good way. I also think our community needs to stop shying away from math. > Focusing on metrics that best measure the behavior we care about is important because benchmarks are essentially an “optimization function” for researchers. I also want to address this, despite it being in the first Re. I will continue to rage against this idea, even if softly put in quotes. No metric is anywhere near close to the behaviors we actually care about. There are no metrics for quality of speech, visual fidelity, vocal realism, and so on. It's impressive that we've done so well when you dig into the metrics we use, but they were selected with care. At the same time, benchmarks are highly limited and especially in discussions of large models (language or vision) these metrics and benchmarks are showing their limitations. Simply due to the alignment of said metrics with desired behavior. We don't desire that the distribution that the LLM learned is indistinguishable from the distribution we used to train it (KL -> 0), but rather we desire that a LLM is able to write language well and perform complex tasks. These do relate, but they are not the same thing. It also makes it disingenuous to compare models with different training sets (such as comparing a JFT pretrained model to something else) due to the nature of what we're actually measuring (e.g. JFT may very well, and likely is, a better approximation of the desired object we're modeling with probability distributions and tuning it to have similar distributional properties to a subset distribution is far easier than training something to learn a distribution in the first place). The desired outcome is, as best we can tell, ineffable. It takes far more than a metric and a benchmark (or several) to quantify the performance of even simple models, let along these beautiful lovecraftian constructs.
- siffland 3y agoI have no idea why this is fun, but on AI chat-bots, i always test 2 + 2(2-2) Should be answer of 2, however on different bots (not just GPT) i have gotten 0, 2, 4 and 6 (all i can understand except the 6). So yeah.......math messes with some bots, who would of guessed.
- nomel 3y agoFor GPT-4, the Wolfram Alpha plugin is great for any maths.
- Sharlin 3y ago"would have" I believe you meant. There's no way even GPT-3.5 fails to solve that ridiculously simple piece of arithmetic. Honestly I'd be surprised if GPT-2 got that wrong. GPT-4 can single-handedly solve vastly more difficult math problems, even though it's handicapped by being merely a language model.
- IanCal 3y ago3.5 & 4 are fine, yes, explaining in steps how to solve it (when prompted purely with "2 + 2(2-2)"). gpt2 completing "2 + 2(2-2) = " returns " 1.5 + 1.5 + 1.5".
- firebaze 3y agoToday I asked ChatGPT4 (openai.com Version) how I can create a copy of a typescript object containing only the fields of a specific type. This is not possible, as typescript (still, unfortunately) doesn't support runtime code generation for its type information. Nevertheless, ChatGPT tried various approaches to "make it work". It reminded me very much of a brainstorming session with Junior Engineers. I couldn't get it to simply state "Not possible". Even after four iterations. Could share the conversation, but would prefer to only share if someone wants me to. TL;DR: LLMs, even the most sophisticated ones, continue to disappoint and underdeliver¹. Posting this as a former ChatGPT4-Enthusiast. I still think we're using LLMs in the wrong way and do not exploit the full potential yet, and there is space (lots of!) to improve, or alternatively, there are limits to accept. ¹ in this case, it's really easy to pinpoint why the answer is wrong because it requires just a little bit of knowledge of the specific technology. But a so-easy-to-spot error, produced by the most expensive, hardest-to-train LLM model we currently have is a telltale sign that something is off. And will continue to be off, even for ChatGPT5.
- nomel 3y ago> ... we're using LLMs in the wrong way ... What do you see as the "right way"?
- firebaze 3y agoTell me. I don't know. Some extremely easy questions will be answered with confidence, but the answers are wrong, while other, hard questions, will be answered correctly. But to tell the difference what is right or wrong, you need to rely on your own judgment or research. In any case, the conversation I referred to today: https://chat.openai.com/share/2dbaeea8-e1dc-4b72-9ae0-01ea5341b8ec https://chat.openai.com/share/2dbaeea8-e1dc-4b72-9ae0-01ea53... As I said, I knew there was no correct answer to begin with aside from "Not possible".
- nomel 3y agoI think the expanse of few-error uses (which could be considered one aspect of "right way") has increased dramatically, with each GPT version, as this article shows. Your example may be trivial in a few years, or with some hyper specific LLM. So, the "right way", to me, is just "what can this LLM do better than the alternatives". I don't think that means we're doing things "wrong" now, but some fact checking peripherals bolted on could definitely help.