12 ms·
GPT-4 as served in the API has been getting 85% on HumanEval (compared to 69.5% claimed here) https://twitter.com/amanrsanger/status/1635751764577361921 https:
by agnokapathetic 3y ago
GPT-4 as served in the API has been getting 85% on HumanEval (compared to 69.5% claimed here)
https://twitter.com/amanrsanger/status/1635751764577361921 https://twitter.com/amanrsanger/status/1635751764577361921
https://github.com/getcursor/eval https://github.com/getcursor/eval
- rushingcreek 3y agoRight, but there's no contamination studies there. I suspect that RLHF data leaked HumanEval into GPT-4. It just seems unlikely to me that GPT-4's coding abilities have improved since March (when 67% was officially reported by OpenAI) given all of the examples and anecdotes about degradation. This is why we use the official numbers.
- refulgentis 3y agoThere weren't any serious examples of degradation. Does only GPT-4 have to suffer a penalty for HumanEval leaking into training data/RLHF data? Ignoring those concerns, it fails a reaonable-ness smell test: We'd have to pretend its the original GPT-4 release from March 2023 until GPT-5 comes out, and only then can OpenAI's work be compared to LLAMA-2 to LLAMA-N.
- rushingcreek 3y agoThere's a couple of things here: 1. I'm not saying we have to wait until GPT-5, we just need an apples-to-apples comparison where contamination is taken into account 2. GPT-4 does not seem to have improved on real-world coding tasks since March, so it's unclear where any purported HumanEval gains could've come from 3. I've personally noticed degradation anecdotally in the GPT-4 June update vs. the original March release
- refulgentis 3y ago1. TL;DR: OpenAI must verify HumanEval data wasn't used in training in order to compare it? 2. Link in the post you replied to. 3. Subjectivity is fine by me! There's a motte & bailey flavor to it if we combine your comment and this one, c.f. "This is why we use the official numbers."
- pclmulqdq 3y agoI think you're assuming that OpenAI is incentivized to benchmark honestly. Like every other company for which a benchmark is a goal, they are not.
- deleted 3y ago[deleted]
- somenameforme 3y agoAlso for a topic like this, subjectivity is all there really is. Even if you create some metric, what you prioritize is going to be subjective. Because performance is going to vary against different sorts of tasks, and there are a literally infinite number of categories of tasks, so it's not like you can ever truly get a fair sampling. Because of this, a sample of subjective opinions is probably much more valuable than any official metric, especially if that metric comes from, as you mentioned, individuals/orgs who are highly motivated to game it endlessly. Even when it comes from an external source you end up with a similar risk of it being gamed. It's like how old school Google puzzle interviews went from seeing who was most clever [in that domain], to seeing who'd booked up the most.
- refulgentis 3y agoWell, no, we have the HumanEval results for the June release.
- somenameforme 3y agoWhich is both (1) a subjective selection to measure the effectiveness of various chatbots and (2) now subject to gaming from companies using opaque/closed/inaccessible/unverifiable systems, like OpenAI.
- lhl 3y ago> 2. GPT-4 does not seem to have improved on real-world coding tasks since March, so it's unclear where any purported HumanEval gains could've come from Once Markdown formatting is accounted for, the June model improves answers on the Leetcode questions from the LLM Drift paper testing to 70% (35/50) vs the March model's 52% (26/50). see: * https://github.com/lchen001/LLMDrift/blob/main/generation/ https://github.com/lchen001/LLMDrift/blob/main/generation/ * https://twitter.com/Si_Boehm/status/1681801371656536068 https://twitter.com/Si_Boehm/status/1681801371656536068
- TuringNYC 3y ago>> given all of the examples and anecdotes about degradation. How many examples and anecdotes about degradation are actually scientific side-by-side studies? I see absurd articles online about ChatGPT usage going down the drain by kids, completely failing to consider even the most basic fact of seasonality and how school is out for the summer!
- devin 3y agoIt takes like 2-3 experiences of receiving a confidently wrong answer to downgrade your usage. If you use a refactoring tool to rename and it misses one, you won’t use it again.
- someplaceguy 3y agoWhile that would likely be my experience with a refactoring tool (unless I didn't have a better alternative), that's not my experience with ChatGPT 4. And that's considering I have very little tolerance for buggy software. There was a period of a few weeks or months in which it seemed like ChatGPT had really degraded to the point of being unusable (although it could have been my biases). However, it seems to be better now (again, my subjective experience). Sometimes I still catch it making really basic mistakes, but most times I can convince it to correct the mistake (especially if I point them out). But what's most amazing to me is how ChatGPT is absolutely brilliant at some things, and not just technical or even obscure topics. Recently, it gave me the most amazing idea for navigating a complex and nuanced social situation I was having difficulty with. And given the constraints of the situation, there was no way I could have gotten that idea otherwise, especially in the allotted time. So despite its flaws and mistakes, I still find it to be a tremendously useful tool, even if only to point me in the right direction.
- ethbr1 3y agoGiven the fact that OpenAI has constant resources (for any given small span of time) and varying demand (users and query type), it's not crazy to think they dynamically adjust to consume all available resources on their side. Obviously the base model would be the same, but aren't there are +/- flavors they could overlay with extra compute? E.g. multi-pass, additional experts, etc. The benefits to giving someone an occasional "magic" answer are too great not to. Have there been any wide studies on same-prompt-different-times?
- mhh__ 3y ago(Chat)GPT-4s practical coding abilities are now 100x because it can code, run the code, and reason about its performance mid-response. They must be using fine tunes for this so the overall model could well be better too
- dontupvoteme 3y agoYou can do that as well, under your complete control. That's a framework they put around the model.
- wordpad25 3y ago"model" is end-to-end, input-to-output, inclusive of the entire framework and it's guardrails and everything else if they are able to detect hallucinations, filter them out and automatically re-run, that's a huge improvement in result, even though core model didn't get new training
- dontupvoteme 3y agoThat's the product. The model is the kernel of the product.
- mhh__ 3y agoThat's what I said.
- oezi 3y agoOnly python though, right?
- ukuina 3y agoIs it possible to learn this power?
- deleted 3y ago[deleted]
- EvgeniyZh 3y agoI have a several arguments why contamination is probably not the main reason of performance difference. When we worked on StarCoder, people ran gpt-4 on MultiPL-E, which doesn't have canonical solutions in the internet, and the performance was higher that what you would expect from official numbers Official contamination analysis shows only minor drop in performance even though contamination is fairly high (you may argue that contamination is higher now or that rlhf has stronger effect) There is significant drop in performance when testing on HumanEval+ [1], which shouldn't happen if model has canonical solutions. BTW why don't you use HumanEval+? [1] https://arxiv.org/abs/2305.01210 https://arxiv.org/abs/2305.01210
- ResearchCode 3y agoThe "intelligence" of large language models needs to be evaluated like the abilities of self-proclaimed psychics. You send your binary to an independent third party and who evaluates it on new problems. It's only a "Human eval" once.
- raincole 3y agoBut the model in OP is fine-tuned by "a proprietary dataset of ~80k high-quality programming problems and solutions". How do we know it's not contaminated by HumanEval too?
- M4v3R 3y agoFrom the OP: > Furthermore, we applied OpenAI's decontamination methodology to our dataset to ensure valid results, and found no contaminated examples.