5 ms·
Great analysis, props to these students for taking the time to challenge such a sensational headline. In the conclusion they mention my biggest problem with the
by underanalyzer 3y ago
Great analysis, props to these students for taking the time to challenge such a sensational headline. In the conclusion they mention my biggest problem with the paper which is that it appears gpt4 grades the answers as well (see section 2.6 "Automatic Grading").
In a way it makes perfect sense that gpt4 can score 100% on a test gpt4 also grades. To be clear the grading gpt4 has the answers so it does have more information but it still might overlook important subtleties in how the real answer differs from the generated answer due to it's own failure to understand the material.
- ghaff 3y agoI noticed that when I read the paper. I know it's hard to scale but I'd want to see competent TAs doing the grading. I also found the distribution of courses a bit odd. Some of it might be just individual samples but intro courses I'd expect to be pretty cookie cutter (for GPT) were fairly far down the list and things I'd expect to be really challenging had relatively good results.
- raunakchowdhuri 3y agoCan attest that the distribution is odd from the test set that we sampled. We've already run the compute to run the zero-shot GPT model on all of the datapoints in the provided test set. We're going through the process now of grading them manually (our whole fraternity is chipping in!) and should have the results out relatively soon. I can say that, so far, it's not looking good for that 90% correct zero-shot claim either.
- mquander 3y agoSince you are here, when I was reading the paper I wondered -- when they show the "zero-shot solve rates", does that mean that they are basically running the same experiment code, but without the prompts that call `few_shot_response` (i.e. they are still trying each question with every expert prefix, and every critique?) It wasn't clear to me at a glance.
- mquander 3y ago> In a way it makes perfect sense that gpt4 can score 100% on a test gpt4 also grades. Even this is overstating it, because for each question, GPT-4 is considered to get it "correct" if, across the (18?) trials with various prompts, it ever produces one single answer that GPT-4 then, for whatever reason, accepts. That's not getting "100%" on a test.
- deleted 3y ago[deleted]
- aeternum 3y agoIn the paper, they at least claimed to manually verify the correct answers.
- sanderjd 3y agoThen - having not read the paper - what is the point of the automated grading?
- mquander 3y agoI just looked again and I didn't see that claim, can you verify? https://arxiv.org/pdf/2306.08997.pdf https://arxiv.org/pdf/2306.08997.pdf If as per the linked critique, some of the questions in the test set were basically nonsense, then clearly they couldn't have manually verified all the answers or they would have noticed that.
- 3y ago
- code51 3y agoThis "GPT4 evaluating LLMs" problem is not limited to this case. I don't know why exactly but everyone seems to have accepted the evaluation of other LLM outputs using GPT4. GPT-4 at this point is being regarded as "ground-truth" with each passing day. Couple this with the reliance on crowd-sourcing to create evaluation datasets and heavy use of GPT3.5 and GPT4 by MTurk workers, you have a big fat feed-forward process benefiting only one party: OpenAI. The Internet we know is dead - this is a fact. I think OpenAI exactly knew how this would play out. Reddit, Twitter and the like are awakening just now - to find that they're basically powerless against this wave of distorted future standards. When sufficiently proven to pass every existing test on Earth, every institution would be so reliant on producing work with GPT that we won't have a "%100 handmade exam" anymore. No problem will be left for GPT to be tackled with.
- YeGoblynQueenne 3y ago>> I don't know why exactly but everyone seems to have accepted the evaluation of other LLM outputs using GPT4. GPT-4 at this point is being regarded as "ground-truth" with each passing day. Why? Because machine learning is not a scientific field. That means anyone can say and do whatever they like and there's no way to tell them that what they're doing is wrong. At this point, machine learning research is like the social sciences: a house of cards, unfalsifiable and unreproducible research built on top of other unfalsifiable and unreproducible research. People simply choose whatever approach they like, cite whatever result they like, because they like the result, not because there's any reason to trust it. Let me not bitch again about the complete lack of anything like objective measures of success in language modelling, in particular. There have been no good metrics, no meaningful benchmarks, for many decades now, in NLP as a whole, but in language generation even more so. This is taught at students in NLP courses (our tutors discussed it in my MSc course) there is scholarship on it, there is a constant chorus of "we have no idea what we're doing" but nothing changes. It's too much hard work to try and find good metrics, build good benchmarks. It's much easier to put a paper on arxiv that shows SOTA results (0.01 more than the best system compared to!). And so the house of cards rises ever towards the sky. Here's a recent paper that points out the sorry state of Natural Language Understanding (NLU) benchmarking: What Will it Take to Fix Benchmarking in Natural Language Understanding? https://aclanthology.org/2021.naacl-main.385/ https://aclanthology.org/2021.naacl-main.385/ There are many more, going back years. There are studies of how top-notch performance on NLU benchmarks is reduced to dust when the statistical regularities that models learn to overfit to in test datasets are removed. Nobody. fucking. cares. You can take your science and go home, we're making billion$$$ here!
- afro88 3y ago> but it still might overlook important subtleties If there's one thing we can be certain of, it's that LLMs often overlooks important subtleties. Can't believe they used GPT4 to also evaluate the results. I mean, we wouldn't trust a student to grade their own exam even when given the right answers to grade with.
- kurthr 3y agoIf people haven't seen it UT Prof Scott Aaronson had GPT4 take his Intro Quantum final exam and had his TA grade it. It made some mistakes, but did surprisingly well with a "B". He even had it argue for a better grade on a problem it did poorly on. Of course this was back in April when you could still get the pure unadulterated GPT4 and they hadn't cut it down with baby laxative for the noobs. https://scottaaronson.blog/?p=7209 https://scottaaronson.blog/?p=7209
- refulgentis 3y agoIt literally did not change. Not one bit. Please, if you're reading this, speak up when people say this. It's a fundamental misunderstanding, there's so much chatter around AI, not much info, and the SnR is getting worse
- kurthr 3y agoWell, I had hoped the sarcastic comparison to cut heroin would make it clear. No, I don't think there's much change at all to GPT-4 (at the API level) and probably not that much at the pre/post language detection and sanitation for apparently psychotic responses.
- jumploops 3y agoI attribute this to two things: 1. People have become more accustomed to the limits of GPT-4, similar to the Google effect. At first they were astounded, now they're starting to see it's limits 2. Enabling Plugins (or even small tweaks to the ChatGPT context like adding today's date) pollute the prompt, giving more directed/deterministic responses The API, as far as I can tell, is exactly the same as it was when I first had access (which has been confirmed by OpenAI folks on Twitter [0]) [0] https://twitter.com/jeffintime/status/1663759913678700544 https://twitter.com/jeffintime/status/1663759913678700544
- JieJie 3y agoIn my experience with Bing Chat, in addition to what you say, there is also some A/B testing going on as well.