5 ms·
GPT-4o hit 54% accuracy on CodeContests with AlphaCodium, up from 48% for GPT-4T
- GavCo 2y agoIt's interesting that with direct prompting there's only a 1 point difference between GPT-4o and GPT-4 turbo, but with the AlphaCodium flow it becomes a substantial 6 point difference. AlphaCodium works by decomposing a competitive programming problem into simple steps and has an automated flow that uses the LLM for each step. It's iterative, so compilation errors and test results are fed back into the model and the model can fix mistakes. IMO this is a much more useful benchmark than a typical eval because it reflects how LLMs are actually used in the real world. It seems like it also surfaces subtle differences in reasoning abilities that zero-shot evals don't capture.
- smt88 2y agoIs a 6-point difference substantial in a statistical sense? To me it seems small enough to be noise.
- lostemptations5 2y agoRight, I would think a 20-30% difference would be significant.
- kibibu 2y agoThat's not what significance means in a statistical sense
- Rastonbury 2y agoI couldn't find the exact test methodology but it would be trivial to rerun it multiple times to get a good average (as well as all other inputs you need to test statistical significance)
- GavCo 2y agovalid point. i'm not a statistician but it does seem to me to be statistically significant. According to the paper, the validation set has 107 problems so it's a difference of 6 problems solved, not just 1-2. Also, it's not a typical eval with multiple-choice answers — these are difficult coding problems from competitions where the answers are strings or integers, so the odds of "guessing" a correct answer are very low.
- aprilthird2021 2y agoEspecially if you consider that the more one writes online about the CodeContests benchmark, the more training data including the CodeContests benchmark data is fed to the next iteration of GPT. We could be doing massive scale over-fitting without realizing it.
- falcor84 2y agoI haven't seen any evidence of over-fitting in the sense of the model getting worse on new examples. Conversely, it seems that we're gradually fitting the models on more and more types of coding questions, such that it's gradually becoming more and more difficult to find well-defined problems that humans can solve easily but AIs cannot.
- wodenokoto 2y agoDepends on how many questions there is. To me it sounds huge, but if there are only 16 questions, then yeah, it’s not significant.
- deleted 2y ago[deleted]